Paper reports Mixture-of-Experts models degrade faster than dense models on repeated training data
- Sources: preprint
- Summary: The preprint compares sparse and dense architectures under repeated training data and reports degradation onset at 4x repetition for Mixture-of-Experts models against 8x for 80M-parameter dense models. The results are the authors' own and are not independently reproduced.
- Why it matters: Data budget planning for a sparse training run cannot reuse the repetition limits measured on dense models.