• Sources: benchmark report, HN discussion
  • Summary: Ben Swerdlow self-published a round-robin benchmark of 171 matches and states that no configuration played beyond beginner level, and that a human beginner running a photon rush would win every game. On the leaderboard Codex Astra at xhigh effort went 18-0, while the three Grok configurations and Claude Haiku each won at most two games. The per-game notes record a multi-agent failure: Codex spawned separate subagents for economy, army production and army control, the subagents did not communicate, and the army subagent sent each finished unit to attack alone rather than massing. The author also reports that models treating a real-time game as turn-based lost while thinking, and says that may explain why some lower-effort settings performed better than higher-effort ones. Method is disclosed, including the round-robin matrix run in parallel on Freestyle VMs with engine data and both harness logs saved per match, and costs are stated as token-based estimates for Codex and Sonnet.
  • Why it matters: A vendor-independent harness reporting that every configuration failed is a usable data point on long-horizon real-time control, and the subagent coordination failure is worth reading before building a similar harness.

send feedback on this story