LLM judges fall short of the noise ceiling on chain-of-thought step importance
- Sources: preprint
- Summary: The preprint asks whether the functional role of a reasoning step is recoverable from the text of that step. It reports LLM judges well below the noise ceiling, with fine-tuned critics still distant from ceiling on correct responses.
- Why it matters: Process reward models and step-level critics assume step text encodes step importance, and this measurement puts that assumption as only partially true.