- Sources: primary, discussion
- Summary: An Amazon Science post dated 2026-08-26 describes a peer-reviewed paper presented at ICML 2026 on aggregating the votes of multiple LLM judges. The method discounts judges whose errors are correlated, and the authors report a 9 to 14 percent improvement over weighted majority vote. The correlation estimate is derived from evaluation logs rather than from new labelled data.
- Why it matters: Teams that add judges to a panel to reduce noise may be adding correlated votes instead, and the method measures that from evaluation logs those teams already collect.
send feedback on this story