- Sources: Charles Azam write-up, HN discussion
- Summary: A blog post compared Claude Fable 5 and GPT-5.6 Sol on a single NP-hard optimization task, with and without a time-boxed
/goal style instruction that directs the agent to work autonomously for a set period. The author framed it as a look at whether the goal instruction improves results on a problem designed for search. - Comments: HN commenters call the top chart confusing because its y-axis is inverted while labeled "lower is better", and argue a single run per model over a large search space is mostly noise rather than a benchmark. One commenter reports the goal-instruction pattern has replaced plan mode in their day-to-day agent workflow.
- Why it matters: Agent "run autonomously toward a goal" instructions are spreading in practitioner workflows, and this thread shows the evidence for them is still anecdotal rather than measured.
send feedback on this story