- Sources: primary
- Summary: The authors argue that spectral allocation accounts for Muon's advantage over Adam and introduce SAMuon. They report SAMuon reaching Muon's validation loss with 13.3 to 24.0 percent fewer training tokens across modded-nanogpt models from 124M to 1B parameters. They state SAMuon carries no persistent optimiser state and no notable extra FLOPs beyond Muon at scale.
- Why it matters: An optimiser with no persistent state and a reported token reduction at fixed loss changes the memory and cost profile of a pretraining run.
send feedback on this story