• Sources: arXiv 2608.13522, submitted 2026-08-13
  • Summary: Vero is a benchmark of repository-level verification tasks, where an agent must produce code and machine-checked proofs across modules rather than for a single function. The paper reports that the strongest coding agent fully solves 27 of the 43 tasks, and that on the hardest repositories the strongest configuration closes no specifications at all.
  • Why it matters: Function-level verified-code benchmarks do not show whether an agent can keep implementation and proof choices coherent across modules, and at repository scale the strongest configuration closes no specifications at all on the hardest repositories.

send feedback on this story