- Sources: arXiv 2607.21273
- Summary: A preprint submitted 2026-07-23 by Yu Wang reports that adding dense prediction rewards to group-relative policy optimization drives language-model agents into what the paper calls a dark room pathology: prediction accuracy converges to 1.0 while task success falls to 0% and episode length pins at the horizon. The paper attributes the failure to GRPO's standard-deviation normalization, reporting that removing only the z-scoring turns the same reward from 0% success back to baseline performance, because all-fail groups combined with the normalization create unbounded pressure that annealing does not remove. Experiments use Qwen3 at 1.7B, 4B, and 8B on ALFWorld, and the paper reports the auxiliary-loss channel gaining roughly 20 points over the reward channel, with a shuffled-label placebo matching true-label performance. The results are a single preprint and not independently reproduced.
- Why it matters: Teams adding shaped rewards to agent RL runs get a named failure mode and a specific normalization term to check before blaming the reward design.
send feedback on this story