• Sources: paper
  • Summary: The paper audits 125 Terminal-Bench tasks that every evaluated model failed and certifies 78 as genuinely unsolved. Of the remaining 47, 14 had broken oracles, 8 were dominated by infrastructure failures, 4 were passable only through verifier bypasses, and 21 have a solvability that the available evidence does not certify either way.
  • Why it matters: Of the remaining 47 tasks, 14 had broken oracles, 8 were dominated by infrastructure failures, 4 were passable only through verifier bypasses, and 21 are uncertified, so an unsolved-task count is weak evidence of model capability.

send feedback on this story