- Sources: primary, discussion
- Summary: The evaluation ran 26 prompt conditions at 80 runs each on a Zstd implementation task, using codex with GPT-5.6 Sol at medium and xhigh effort levels. No condition wildly outperformed, and the default condition, which gave the agent no testing instruction, scored well above the average across conditions. Agents given the name of a technique or a library applied it superficially rather than to the code most likely to hold the bug.
- Why it matters: Prompt boilerplate that names fuzzing or property testing did not outperform the default on this task and this model.
send feedback on this story