• Sources: primary, discussion
  • Summary: The post applies two measures of what it calls code sloppiness, verbosity and erosion, and states both were introduced to it by the SlopCodeBench paper rather than defined by the post itself. It reports verbosity of 0.15 plus or minus 0.06 for established repositories against 0.33 plus or minus 0.10 for agent-written code, and erosion of 0.31 plus or minus 0.17 against 0.68 plus or minus 0.20. SlopCodeBench is a multi-round benchmark in which context is erased between checkpoints and every checkpoint's test suite must pass.
  • Why it matters: Under SlopCodeBench, where context is erased between checkpoints and every checkpoint's tests must pass, the post reports a strict pass rate of zero for the models it tested, which a footnote states include GPT 5.6 and not Fable 5.1 or Astra.

send feedback on this story