- Sources: primary, discussion
- Summary: The paper organizes procedural knowledge into procedure-relation-procedure triplets and localizes the agent's active node to bias the next action. An LLM refiner contrasts failed and successful trajectories and commits graph edits that hold up on held-out validation. The preprint was submitted 2026-09-08 and is not independently reproduced.
- Why it matters: It targets the failure mode where long trajectories lose the objective and repeat unproductive tool calls.
- Follow-up: Independent reproduction, and whether the learned graph edits transfer across task families.
send feedback on this story