- Sources: primary, discussion
- Summary: Dan Abramov, who states he is not a mathematician, describes running Codex sessions in fixed roles, a project manager that commits work, two math agents, a red agent, a random agent and a Lean agent, with a relay agent emulating a group chat, and resetting sessions when they drifted. He reports that one-shot prompting produced invented terminology he could not pass to a mathematician, that a claimed complete proof was withdrawn after a fresh session found a circular step which also invalidated earlier results, and that the fix was refusing to build on unformalized claims and grounding the work by having mathematicians confirm typo fixes in a peer-reviewed reference. The resulting Lean proof passed the Palomar registry mechanical checks, is not independently verified by mathematicians, and the author invites refutation, so the mathematical claim stays developing.
- Why it matters: The account is the most detailed public description of a long-running multi-agent harness and names its failure modes, including a proof that read as complete and was circular.
- Follow-up: Track independent mathematician verification of the Lean proof, which advances the open follow-up on AI-credited proofs.
send feedback on this story