- Sources: paper
- Summary: The preprint compares seven methods against repeated sampling at matched token cost, counting every critique, reflection, debate and checking token, across open models of 1.5B, 3B and 7B parameters on two mathematics benchmarks at 150 questions each. That design gives 36 comparisons, each paired by question, and the paper reports that no method beats repeated sampling at equal cost, that ten of the 36 are reliably worse, all of them methods where the model inspects its own output, and that all 18 self-inspection comparisons are negative. The two kinds of self-inspection separate as models grow: the Best-of-N choosing gap falls from 8.0 and 11.3 points at 1.5B to 2.0 and 1.3 at 7B, which the paper describes as no longer distinguishable from zero, while Self-Refine and a forced Reflexion stay 3.6 to 10.1 points below baseline at 7B.
- Why it matters: Self-critique and reflection loops are a common agent design pattern, and this measurement bounds them only for open models up to 7B, where the penalty for having the model choose among its own samples closes with scale but the penalty for having it rewrite its own output does not.
- Follow-up: Single-author v1 preprint with no peer review. Track whether the result holds at larger model scale.
send feedback on this story