• Sources: primary
  • Summary: The benchmark separates whether a migration actually happened from whether behaviour survived, because an agent can copy the original implementation to pass the tests, which the authors call Blindness. Across 8 frontier models and 26 model-effort configurations, 28 of 520 runs pass all three stages, 5.4 percent, and 13 of the 20 tasks receive no accepted solution. Agents score 31.4 out of 100 on build toolchain rewrites against 5.6 on language rewrites, and the best model scores 47.0 out of 100 overall. The paper is a preprint submitted 2026-08-24 with no peer review and no independent replication, so both the protocol and the scores are unreviewed.
  • Why it matters: The gap between toolchain rewrites and language rewrites is the number to use when scoping an agent-led migration.

send feedback on this story