• Sources: preprint
  • Summary: The preprint compares sparse and dense architectures under repeated training data and reports degradation onset at 4x repetition for Mixture-of-Experts models against 8x for 80M-parameter dense models. The results are the authors' own and are not independently reproduced.
  • Why it matters: Data budget planning for a sparse training run cannot reuse the repetition limits measured on dense models.

send feedback on this story