- Sources: preprint
- Summary: The private, contamination-controlled suite holds 256 repository and post-cutoff contest tasks, and each paired same-model contrast ran the same 80 of them under a vendor harness and under a neutral one, reporting -1.25 points for Opus 4.8 between claude-agent-sdk and deepagents and +1.25 for GPT-5.5 between the openai-codex SDK and deepagents, both with confidence intervals spanning zero. The Opus average hides opposite strata, with the native harness trailing by 9.0 points on repository tasks and leading by 23.7 on contest tasks, and the author states that partition was chosen after seeing the data and needs a designed replication. The tasks stay private and the results are unreproduced by third parties.
- Why it matters: One finding is directly operational: 22 of 81 runs cancelled at the wall-clock ceiling had already produced a passing patch, so completion and correctness are not the same measurement.
send feedback on this story