- Sources: HumanLayer write-up, arXiv 2603.24755, HN 49076391
- Summary: A write-up in the HumanLayer advanced-context-engineering repository reports running SlopCodeBench against three models in one session of about six hours, covering 3 of the benchmark's 36 problems across 17 checkpoints. The reported result is Opus 5 at 4 of 17 strict passes, against 1 of 17 each for Opus 4.8 and Sonnet 5. Opus 5 produced 29,065 source lines against roughly 9,000 for each of the other two, with 51 percent of that output tests. None of the three models finished any of the three problems clean. The evidence base is one practitioner and one run, and the author hedges the result himself. The benchmark is arXiv 2603.24755, submitted 2026-03-25 and revised 2026-05-07, which defines 36 problems across 196 checkpoints. Its own abstract reports that across 15 coding agents no agent fully solves any problem end to end and the best agent passes 14.8 percent of checkpoints, so the practitioner run is not comparable to a paper leaderboard position.
- Comments: HN commenter killingtime74 asks why GPT-5.6, GLM 5.1, and Kimi K3 were left out and offers to run them. Other replies report dissatisfaction with Opus 5 on their own work, which is opinion rather than measurement.
- Why it matters: SlopCodeBench withholds later requirements instead of stating the whole problem up front, so it measures whether a coding agent keeps a codebase changeable across checkpoints rather than whether it solves one stated task, which is the property a team is betting on when it runs agents unattended.
- Follow-up: Watch for a run that covers GPT-5.6, GLM 5.1, and Kimi K3.
send feedback on this story