How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization
Abstract
Recent frontier large language models predominantly rely on Mixture-of-Experts (MoE) architectures. Despite empirical progress, there is still no principled understanding of how hyperparameters should scale with network width , expert width , number of experts , sparsity , and depth to ensure both stability and optimal performance at scale. We take a principled step toward resolving this gap by analyzing three different scaling regimes: (I) co-scaling , (II) co-scaling , and (III) full proportional scaling of , and . For each regime, we develop a novel Dynamical Mean Field Theory (DMFT) description of the limiting training dynamics of MoEs that provides a formal foundation for our analysis. Within this framework, we derive the unique parameterization for SGD and Adam satisfying all maximal-update () desiderata. We then show that the resulting P prescription does not reliably induce monotonic improvement with scale or robust learning-rate transfer. We trace these pathologies to scale-dependent observables in the aggregation dynamics, which motivates a refined set of desiderata that we term maximal scale stability. Guided by this principle, we derive a Maximally Scale-Stable Parameterization (MSSP) for both SGD and Adam in all three scaling regimes, and characterize the corresponding limiting dynamics - qualitatively distinct from the P limit - through a separate DMFT analysis. Experiments verify that MSSP robustly recovers learning rate transfer and monotonic improvement with scale across regimes. Combined with existing depth-scaling theory, these results provide a complete scaling prescription for MoE architectures as a function of width, depth, expert width, and number of experts.
Cite
@article{arxiv.2605.14200,
title = {How to Scale Mixture-of-Experts: From muP to the Maximally Scale-Stable Parameterization},
author = {Leena Chennuru Vankadara and Moritz Haas and Luke Hayward and Sebastian Bordt and Alessandro Breccia},
journal= {arXiv preprint arXiv:2605.14200},
year = {2026}
}