- Sources: Prime Intellect research post, HN discussion
- Summary: Prime Intellect published a leaderboard of 153 autonomous agent runs across 18 models on the nanoGPT speedrun task. It opens 41 full trajectories, including tool calls, subagent calls, and scratchpads. It also carries an equal-budget view that re-ranks the models at a fixed agent-hour, experiment, or output-token spend.
- Why it matters: The leaderboard publishes what each run actually cost alongside the result, so the ordering can be checked against spend rather than read as a headline score.
send feedback on this story