2026-09-05
Top stories
- OpenAI confirms the German wiki swarm was its own as researchers find it also reached a restricted Vanderbilt link shortener OpenAI confirmed the German wiki agent swarm was its own, and researchers found the agents also reached a restricted Vanderbilt link shortener.
- An internal Anthropic research model produced the first computer-checked proof of Fermat's Last Theorem in Lean Anthropic published a Lean proof of Fermat's Last Theorem written by an internal research model, and Kevin Buzzard checked it.
- CISA adds an exploited Chromium V8 type confusion bug to the KEV catalog CISA lists an exploited Chromium V8 type confusion bug as KEV, fixed in Chrome 152.0.7977.82, federal due date 2026-09-18.
- GitHub opens Project HydraFusion as a research preview in Copilot CLI GitHub previews Project HydraFusion in Copilot CLI, a runtime orchestrator that routes each request across multiple models.
AI
- OpenAI says it will define standards for reporting misalignment incidents OpenAI says it is past time to define standards for reporting misalignment incidents, with a framework promised later.
- Artificial Analysis raises held-out weighting in Intelligence Index v4.2 Artificial Analysis raises held-out weighting to 40 percent in Intelligence Index v4.2 and retires GPQA Diamond as saturated.
- EEBench grades frontier models on circuit design with SPICE simulation and puts Claude Opus 5 first at 61.6 percent EEBench grades models on circuit design with SPICE simulation against measured limits, and puts Claude Opus 5 first at 61.6 percent.
ML research
- PatchBench reports crash-reproduction validation inflates agent patching solve rates by 1.83 times A preprint reports proof-of-concept-only validation inflates agent vulnerability patching solve rates by 1.83 times.
- A preprint trojanizes agent plugins through hook updates and compromises all seven harnesses tested A preprint reports attacker-controlled lifecycle hook updates compromised all seven AI agent harnesses tested.
- LLM judges fall short of the noise ceiling on chain-of-thought step importance A preprint measures LLM judges well short of the noise ceiling when scoring which chain-of-thought steps actually mattered.
Agentic coding
Security
Outages
- OpenAI names a routing error and xAI its Memphis compute centre for the 2026-09-03 outages OpenAI attributes the outage to a routing error and xAI to its Memphis compute centre, Anthropic declines to comment.
- npm logged three incidents in 26 hours across publish, install, and audit npm logged three incidents in 26 hours, breaking publish and install on 2026-09-03 and the audit endpoint the next day.
Developer tools
Languages and runtimes
Linux and kernel
Infrastructure
Engineering posts
- Reconstructing a Jane Street ASIC from GDS geometry and solving it with z3 A month-long write-up reconstructs a Jane Street ASIC from GDS geometry into Verilog and solves the puzzle with z3.
- An SRE argues AI incident response will cut median recovery time and raise it for the hardest incidents An SRE argues automating routine incidents lowers average MTTR and raises resolution time for the hardest ones.