• Sources: preprint
  • Summary: The preprint asks whether the functional role of a reasoning step is recoverable from the text of that step. It reports LLM judges well below the noise ceiling, with fine-tuned critics still distant from ceiling on correct responses.
  • Why it matters: Process reward models and step-level critics assume step text encodes step importance, and this measurement puts that assumption as only partially true.

send feedback on this story