- Sources: primary
- Summary: The paper reports an audit of hosted LLM judges across 52,988 request attempts, measuring whether rankings produced through shared endpoints hold across repeated runs. The authors registered reliability thresholds before collection and report that the measured rankings did not meet them.
- Why it matters: Anyone gating training data or a leaderboard on a hosted judge is treating a model name as a fixed instrument, and this audit measures how far that assumption failed.
- Follow-up: Watch for peer review, and for self-hosted replication under sustained load, the condition the paper reports it did not clear.
send feedback on this story