English

$\mu$-Parametrization for Mixture of Experts

Machine Learning 2025-10-10 v2

Abstract

Recent years have seen a growing interest and adoption of LLMs, with Mixture-of-Experts (MoE) emerging as a leading architecture in extremely large models. Currently, the largest open-source models reach over 11T parameters. At such scales, hyperparameter tuning becomes prohibitively expensive. Precisely for this reason, the μ\muTransfer is becoming a key technique. It allows for seamless transfer of optimal hyperparameters across model scales, resulting in a huge reduction in tuning costs. However, existing work has primarily focused on dense LLMs, leaving MoE architectures unexplored. In this work, we derive a μ\mu-Parameterization for MoE, providing theoretical guarantees for feature learning across model widths. Our experiments demonstrate that the optimal learning rate reliably transfers across model sizes, establishing a foundation for efficient hyperparameter tuning in large-scale MoE models.

Keywords

Cite

@article{arxiv.2508.09752,
  title  = {$\mu$-Parametrization for Mixture of Experts},
  author = {Jan Małaśnicki and Kamil Ciebiera and Mateusz Boruń and Maciej Pióro and Jan Ludziejewski and Maciej Stefaniak and Michał Krutul and Sebastian Jaszczur and Marek Cygan and Kamil Adamczewski and Jakub Krajewski},
  journal= {arXiv preprint arXiv:2508.09752},
  year   = {2025}
}
R2 v1 2026-07-01T04:48:02.609Z