- Sources: paper, discussion
- Summary: The authors report that quantized reasoning models reached the correct answer in intermediate steps but failed to emit it in up to 52 percent of failures, spending additional chain-of-thought tokens instead. A training-free logit penalty on overthinking markers reduces chain-of-thought length by 12 to 23 percent while preserving or improving accuracy across 5 models, 3 quantization methods and 5 benchmarks, and reduces overthinking errors by up to 58 percent. Submitted 2026-05-29.
- Why it matters: Anyone quantizing a reasoning model to cut serving cost may be paying that saving back in chain-of-thought tokens.
send feedback on this story