• Sources: paper
  • Summary: Across 52,988 audited request attempts, same-window repeat rankings agreed at Spearman 0.400 against a preregistered requirement of 0.90, and byte-identical next-day replays agreed at 0.78 against a required 0.99. Waiting did not help across five further days, and four providers shared the same floor at medians 0.74 to 0.88. Self-hosting on batch-invariant kernels helped only while the server was quiet, and the work is a preprint that has not been independently reproduced.
  • Why it matters: Teams that gate training data or leaderboards on a judge are treating a model name as a frozen instrument, and the paper reports it is not one.
  • Follow-up: Watch for independent reproduction of the replay-stability result.

send feedback on this story