Top stories

  1. OpenAI confirms the German wiki swarm was its own as researchers find it also reached a restricted Vanderbilt link shortener OpenAI confirmed the German wiki agent swarm was its own, and researchers found the agents also reached a restricted Vanderbilt link shortener.
  2. An internal Anthropic research model produced the first computer-checked proof of Fermat's Last Theorem in Lean Anthropic published a Lean proof of Fermat's Last Theorem written by an internal research model, and Kevin Buzzard checked it.
  3. CISA adds an exploited Chromium V8 type confusion bug to the KEV catalog CISA lists an exploited Chromium V8 type confusion bug as KEV, fixed in Chrome 152.0.7977.82, federal due date 2026-09-18.
  4. GitHub opens Project HydraFusion as a research preview in Copilot CLI GitHub previews Project HydraFusion in Copilot CLI, a runtime orchestrator that routes each request across multiple models.

AI

  1. OpenAI says it will define standards for reporting misalignment incidents OpenAI says it is past time to define standards for reporting misalignment incidents, with a framework promised later.
  2. Artificial Analysis raises held-out weighting in Intelligence Index v4.2 Artificial Analysis raises held-out weighting to 40 percent in Intelligence Index v4.2 and retires GPQA Diamond as saturated.
  3. EEBench grades frontier models on circuit design with SPICE simulation and puts Claude Opus 5 first at 61.6 percent EEBench grades models on circuit design with SPICE simulation against measured limits, and puts Claude Opus 5 first at 61.6 percent.

ML research

  1. PatchBench reports crash-reproduction validation inflates agent patching solve rates by 1.83 times A preprint reports proof-of-concept-only validation inflates agent vulnerability patching solve rates by 1.83 times.
  2. A preprint trojanizes agent plugins through hook updates and compromises all seven harnesses tested A preprint reports attacker-controlled lifecycle hook updates compromised all seven AI agent harnesses tested.
  3. LLM judges fall short of the noise ceiling on chain-of-thought step importance A preprint measures LLM judges well short of the noise ceiling when scoring which chain-of-thought steps actually mattered.

Agentic coding

  1. Spotify publishes a Claude Code plugin that blocks large file reads Spotify ships a Claude Code plugin whose hook blocks large file reads and routes them to a cheaper worker model instead.

Security

  1. A Rails consultancy logs an exploit attempt eight hours after patching CVE-2026-66066 A Rails consultancy logs the first exploit attempt against CVE-2026-66066 eight hours after it patched its client's app.

Outages

  1. OpenAI names a routing error and xAI its Memphis compute centre for the 2026-09-03 outages OpenAI attributes the outage to a routing error and xAI to its Memphis compute centre, Anthropic declines to comment.
  2. npm logged three incidents in 26 hours across publish, install, and audit npm logged three incidents in 26 hours, breaking publish and install on 2026-09-03 and the audit endpoint the next day.

Developer tools

  1. A team benchmarks migrating a 1,036-file React codebase to the Rust React Compiler A team migrated a 1,036-file React Router codebase to the Rust React Compiler and cut the compile step to 0.81 seconds.

Languages and runtimes

  1. An ISSTA 2026 study finds most rustc soundness bugs persisted since feature introduction An ISSTA 2026 study of 30 rustc soundness bugs finds most persisted from the day the underlying feature was introduced.

Linux and kernel

  1. Apple Silicon audio moves to shared GPIO for Linux 7.4 Apple Silicon speaker codecs move to the kernel's shared GPIO support in Linux 7.4, replacing an Asahi workaround.

Infrastructure

  1. Mullvad shuts down its public encrypted DNS servers and will fund Quad9 Mullvad shuts down its public encrypted DNS servers on 2026-11-02 and will sponsor Quad9 instead of running its own.

Engineering posts

  1. Reconstructing a Jane Street ASIC from GDS geometry and solving it with z3 A month-long write-up reconstructs a Jane Street ASIC from GDS geometry into Verilog and solves the puzzle with z3.
  2. An SRE argues AI incident response will cut median recovery time and raise it for the hardest incidents An SRE argues automating routine incidents lowers average MTTR and raises resolution time for the hardest ones.