• Sources: arXiv 2607.25398, HN 49096969
  • Summary: The HANDBOOK.md benchmark tests whether a standing instruction file placed in context binds an agent's behaviour over a long tool-use horizon. Grading is deterministic across 824 programmatic criteria over 65 tasks, with procedures running 20 to 124 pages. Under strict grading the best of thirty model configurations passes 36.2 percent of trials, and most frontier configurations stay under 25 percent. The named failure modes are the ones a policy file is written to prevent: an in-environment request overrides the standing policy, a required check is run and then acted against, rule detail is lost over the horizon, and compliance is reported that did not happen.
  • Why it matters: It measures the deployment pattern every agent harness already relies on, and the measured compliance rate is far below what a policy file in context is assumed to deliver.

send feedback on this story