• Sources: primary
  • Summary: The authors fix an order on the three sources of training nondeterminism, namely GPU kernel reductions, batch ordering across a data-parallel cluster, and inter-node and intra-node collective communication, so any single step of a large distributed run can be replayed on one commodity machine and checked bitwise against the published trajectory. Because replaying a whole run on one machine is infeasible, they propose collective verification in which independent auditors each certify individual steps. They release Open-1B with its full pretraining dataset, every intermediate checkpoint, the training code and the audit harness.
  • Why it matters: Published weights and recipes do not prove a checkpoint came from the declared recipe, which is the gap that leaves room for undisclosed data or a backdoor, and bitwise step replay is a check a third party can run on one machine.
  • Follow-up: The preprint is not peer reviewed and the reproducibility claim has not been independently exercised, so track a third-party replay of a published step.

send feedback on this story