• Sources: primary, discussion, reddit
  • Summary: The release post states all tasks now use a flat 8-hour agent timeout, that 8 tasks were removed for saturation, refusals, public solutions, and unresolved quality or platform issues at two each, and that 19 tasks were fixed. It defines saturated as solved 5 of 5 times by every class within every family of the latest model generation, and states the project is now a continuous benchmark on semantic versioning rather than releasing sequels, running on the Harbor framework. It reports Sonnet 5 using 21.6 billion tokens on its leaderboard run against 6.5 billion for Opus 5, with large variance in agent execution time, and discloses that leaderboard experiments are supported by grants from OpenAI, Anthropic, Z.ai, SpaceX AI, and the Laude Institute.
  • Why it matters: Removing tasks and changing agent resources are breaking changes that require re-running trials, so a 3.0 number and a 4.0 number are not the same measurement, and anyone quoting an agent leaderboard has to state which version produced it.

send feedback on this story