• Sources: primary
  • Summary: A preprint posted 2026-09-03 argues that repository-level coding benchmarks measure whether a patch passes functional tests and ignore the review-derived constraints that decide whether a patch is accepted. SWE-Gate builds 303 repository-level repair instances across 75 open-source Python repositories, deriving constraints from real pull request review comments and shipping separate functional and constraint tests plus non-compliant and gold patches for each instance. Across four LLM backends under one common agent scaffold, 221 of 644 repairs that passed the functional tests failed the review constraints.
  • Why it matters: A pass rate quoted for a coding agent describes functional tests only, and this measurement puts the gap to the full specification at about a third.
  • Follow-up: Watch for peer review and for independent reproduction from the published replication package.

send feedback on this story