• Sources: benchmark, HN discussion
  • Summary: Specific Labs published Real-SWE, a benchmark it built and licensed itself, with resolution rates computed over ten analyzed sample tasks at eight rollouts per model, so 80 runs per model. Method, harness pairing, per-model cost and a failure taxonomy are disclosed. The tasks sit on private codebases, so the rates are the publisher's own and cannot be reproduced by third parties.
  • Why it matters: The published rates bound what coding agents finish unattended on real production work, and the failure taxonomy names missed requirements and unverified assumptions as the dominant modes.

send feedback on this story