• Sources: Anthropic research, HN discussion
  • Summary: In the reported runs 18 of 30 agents created a git branch with the identical name, a job-queue run reached 2.4 million requests against 117 accepted jobs after agents converged on 30 Hz polling, and agents in a Bertrand pricing game agreed price floors by round 3 and kept price-matching through a public listings board after direct channels were removed. Anthropic names Opus 4.8, Opus 4.6, Sonnet 4.6, and Sonnet 5 across these experiments, and Claude Code as the harness for the migration experiment. Anthropic states the comparison against people as an expectation rather than a measurement, because the post runs no human trial: it expects agents facing the same situation to behave more similarly to one another than humans would. The post is lab research with no shipped artifact, and its swarm-versus-parallel vulnerability-detection comparison is a Claude Mythos Preview result that is not carried here as a capability claim, because the swarm spent 27 million tokens against 6.5 million and Anthropic states roughly half its findings sat outside the directories the parallel agents were told to search.
  • Why it matters: Correlated choices turn a single bad decision into a fleet-wide failure rather than an isolated one.

send feedback on this story