English

Practical Efficiency of Muon for Pretraining

Machine Learning 2025-05-21 v4 Machine Learning

Abstract

We demonstrate that Muon, the simplest instantiation of a second-order optimizer, explicitly expands the Pareto frontier over AdamW on the compute-time tradeoff. We find that Muon is more effective than AdamW in retaining data efficiency at large batch sizes, far beyond the so-called critical batch size, while remaining computationally efficient, thus enabling more economical training. We study the combination of Muon and the maximal update parameterization (muP) for efficient hyperparameter transfer and present a simple telescoping algorithm that accounts for all sources of error in muP while introducing only a modest overhead in resources. We validate our findings through extensive experiments with model sizes up to four billion parameters and ablations on the data distribution and architecture.

Keywords

Cite

@article{arxiv.2505.02222,
  title  = {Practical Efficiency of Muon for Pretraining},
  author = {Essential AI and : and Ishaan Shah and Anthony M. Polloreno and Karl Stratos and Philip Monk and Adarsh Chaluvaraju and Andrew Hojel and Andrew Ma and Anil Thomas and Ashish Tanwer and Darsh J Shah and Khoi Nguyen and Kurt Smith and Michael Callahan and Michael Pust and Mohit Parmar and Peter Rushton and Platon Mazarakis and Ritvik Kapila and Saurabh Srivastava and Somanshu Singla and Tim Romanski and Yash Vanjani and Ashish Vaswani},
  journal= {arXiv preprint arXiv:2505.02222},
  year   = {2025}
}
R2 v1 2026-06-28T23:20:48.298Z