• Sources: primary
  • Summary: The paper, whose arXiv comments field states it is the camera-ready version accepted at ADMA 2026, the International Conference on Advanced Data Mining and Applications, in its Special Session on Responsible Data Intelligence, audits 254 SWE-bench submissions across four splits without rerunning any model. On Verified the leading two entries each resolve 396 of 500 instances, the top ten share 285 successes and 51 failures, and only 164 instances distinguish their outcomes at all. Exact paired McNemar tests separate none of the 29 adjacent pairs in the Verified top thirty at alpha 0.05, while the larger Test split separates 14 of 23, and observed within-model scaffold ranges reach 29.8 percentage points against an 8.8-point spread across the entire top thirty.
  • Why it matters: Picking a coding agent on a one-point leaderboard gap is not supported by the published verdicts, because the harness moves the score more than the model does.
  • Follow-up: Track whether anyone outside the author group adopts the released partition and five-step audit protocol when reporting a SWE-bench result.

send feedback on this story