• Sources: primary, discussion
  • Summary: The paper fixes the execution loop and varies planning, action space and context management across 176 matched settings on SWE-Bench Verified and Terminal-Bench 2.1 with four models. It reports that staging rule-based elision before LLM summarization gives the strongest overall efficiency among the context-management strategies, that making elided content recoverable adds machinery models rarely use and gains no accuracy, that planning shifts from an accuracy scaffold on weaker models to a cost saver on stronger ones, and that predefined tools improve performance for models with weaker bash proficiency while bash-capable models run a bash-only interface at substantially lower cost, especially on command-line-centric tasks.
  • Why it matters: The matched-setting comparison is the one most harness write-ups skip, so the results are directly actionable for anyone choosing an action space or a context policy.

send feedback on this story