• Sources: primary, discussion
  • Summary: The study, dated 2026-08-12, runs two to three rollouts per cell across three Claude models and is labelled by its author as a preliminary pilot. Attaching source lines to the same semantic results cut follow-up file reads from 15.2 to 3.2 per episode and lifted pass@1 from 0.67 to 0.83. The author discloses a product interest in AgentConnect.
  • Why it matters: What a retrieval tool returns is a design decision separate from what it retrieves, and the measured gain came from the output shape rather than from better matches.
  • Follow-up: Watch for a run with enough rollouts to separate the effect from noise.

send feedback on this story