中文

Muon 运动:基于正则化 Muon 优化器的哈密顿概率梯度流视角

机器学习 2026-05-25 v1 机器学习 统计理论 统计理论

摘要

我们在矩阵参数空间上定义的概率测度上构建了一个梯度流,由正则化 Muon 引起,这是理想化 Muon 优化器的一个解析平滑版本。关键观察是,正则化正交化映射是平滑凯恩-斯塔克偶极化的梯度。这将(正则化)Muon 更新识别为更新变量中的镜像/近似步骤,动量则作为对偶坐标起作用。我们利用这一结构,将 Muon 从单一矩阵参数提升至形式为 J(ρ)=R(Fdρ)J(\rho)=R\left(\int F d \rho\right) 的有限粒子概率目标,其中动力学来源于神经网络训练中的平均场描述,并推导了惯性连续时间极限。通过这一结构,我们推导了在步长和动量比例缩放下的有限粒子连续时间极限,然后转化为参数-动量对概率分布上的相空间均匀方程。 resulting flow can be shown to be a damped Hamiltonian probability dynamics whose kinetic energy is induced by the regularized Muon mirror potential. We prove an exact Hamiltonian dissipation identity, showing that the Hamiltonian energy decreases monotonically. While the target objective itself need not be monotone along the inertial Muon dynamics, under additional gradient-dominance, bounded-momentum, and curvature/alignment assumptions, we obtain continuous and discrete-time exponential convergence rates for the objective gap. We also study the well-posedness of the mean-field limit equation and establish propagation of chaos guarantees for the interacting particle system. Finally, we extend the formulation to Hilbert-valued feature maps on product matrix spaces, yielding a blockwise Muon probability flow applicable to smooth transformer mixture-of-experts models.

关键词

引用

@article{arxiv.2605.23871,
  title  = {Move on Muon : A Hamiltonian probability gradient flow perspective of Muon optimizer},
  author = {Aratrika Mustafi and Soumya Mukherjee and Bharath K. Sriperumbudur},
  journal= {arXiv preprint arXiv:2605.23871},
  year   = {2026}
}