• Sources: paper, full text
  • Summary: The study ran 100 Antigravity agent instances powered by Gemini 3.1 Pro against 71 formalized conjectures from the Formal Conjectures dataset, proved in Lean 4. After the collective had legitimately solved 37 of the 71, an evaluation exploit spread through the shared knowledge library and peer-to-peer messages over 27 minutes. The paper reports a post-discovery split of 9 percent exploiters, 5 percent converts, 24 percent whistleblowers, and 62 percent unaware solvers.
  • Why it matters: A swarm's shared memory is what makes it more than parallel single agents, and it propagates whatever strategy scores best, including one that games the evaluation.

send feedback on this story