English
Related papers

Related papers: Muown: Row-Norm Control for Muon Optimization

200 papers

The Multi-Objective Evolutionary Algorithm based on Decomposition (MOEA/D) is a popular algorithm for solving Multi-Objective Problems (MOPs). The main component of MOEA/D is to decompose a MOP into easier sub-problems using a set of weight…

Neural and Evolutionary Computing · Computer Science 2021-09-14 Yuri Lavinas , Abe Mitsu Teru , Yuta Kobayashi , Claus Aranha

Large scale neutrino detectors and muon tomography rely on the muon direction in the detector to infer the muon's or parent neutrino's origin. However, muons accumulate deflections along their propagation path prior to entering the…

Instrumentation and Methods for Astrophysics · Physics 2023-02-03 Pascal Gutjahr , Jean-Marco Alameddine , Alexander Sandrock , Jan Soedingrekso , Mirco Hünnefeld , Wolfgang Rhode

We extend the previously published model that distinguishes between the diffusive motion of diamagnetic muons and the dynamics of ions around the muon in matter, and propose a generalized model for {\sl paramagnetic muons} (Mu$^0$s, bound…

Atomic Physics · Physics 2025-05-28 Ryosuke Kadono , Takashi U. Ito

We introduce Pion, a spectrum-preserving optimizer for large language model (LLM) training based on orthogonal equivalence transformation. Unlike additive optimizers such as Adam and Muon, Pion updates each weight matrix through left and…

Machine Learning · Computer Science 2026-05-13 Kexuan Shi , Hanxuan Li , Zeju Qiu , Yandong Wen , Simon Buchholz , Weiyang Liu

Scaling neural network training increasingly depends on synchronous data-parallelism, yet full-precision gradient all-reduce imposes a severe communication bottleneck. We propose Decoupled Momentum Optimization (DeMo), a drop-in replacement…

Machine Learning · Computer Science 2026-02-10 Bowen Peng , Lizhang Chen , Baiyu Su , Jeffrey Quesnelle , Diederik P. Kingma , Qiang Liu

Training a modern machine learning architecture on a new task requires extensive learning-rate tuning, which comes at a high computational cost. Here we develop new Polyak-type adaptive learning rates that can be used on top of any momentum…

Machine Learning · Computer Science 2024-06-06 Fabian Schaipp , Ruben Ohana , Michael Eickenberg , Aaron Defazio , Robert M. Gower

Equivariant neural networks encode geometric symmetries by construction, yet they are often difficult to optimize and can underperform less constrained architectures. A growing body of work addresses this through architectural modifications…

Machine Learning · Computer Science 2026-05-28 Teodor-Mihai Stupariu , Andrei Manolache

Cosmic-ray muon sources exhibit distinct scattering angle distributions when interacting with materials of different atomic numbers (Z values), facilitating the identification of various Z-class materials, particularly those radioactive…

Instrumentation and Detectors · Physics 2025-04-18 Haochen Wang , Zhao Zhang , Pei Yu , Yuxin Bao , Jiajia Zhai , Yu Xu , Li Deng , Sa Xiao , Xueheng Zhang , Yuhong Yu , Weibo He , Liangwen Chen , Yu Zhang , Lei Yang , Zhiyu Sun

The TWIST collaboration has performed new measurements of two of the parameters that describe muon decay: $\rho$, which governs the shape of the overall momentum spectrum, and $\delta$, which governs the momentum dependence of the…

We propose adaptive weight decay, which automatically tunes the hyper-parameter for weight decay during each training iteration. For classification problems, we propose changing the value of the weight decay hyper-parameter on the fly based…

Machine Learning · Computer Science 2023-12-05 Amin Ghiasi , Ali Shafahi , Reza Ardekani

Large language models deliver strong reasoning and tool-use skills, yet their computational demands make them impractical for edge or cost-sensitive deployments. We present \textbf{Xmodel-2.5}, a 1.3-billion-parameter small language model…

Machine Learning · Computer Science 2025-11-26 Yang Liu , Xiaolong Zhong , Ling Jiang

Empirical scaling laws prescribe how to allocate parameters, data, and compute, while maximal-update parameterization ($\mu$P) enables learning-rate transfer across widths by equalizing early-time update magnitudes. However, in modern…

Machine Learning · Computer Science 2025-10-20 Zhiyuan Fan , Yifeng Liu , Qingyue Zhao , Angela Yuan , Quanquan Gu

Over-parameterized deep networks trained using gradient-based optimizers are a popular choice for solving classification and ranking problems. Without appropriately tuned $\ell_2$ regularization or weight decay, such networks have the…

Machine Learning · Computer Science 2021-08-13 Aman Gupta , Rohan Ramanath , Jun Shi , Anika Ramachandran , Sirou Zhou , Mingzhou Zhou , S. Sathiya Keerthi

Scaling large models requires optimization strategies that ensure rapid convergence grounded in stability. Maximal Update Parametrization ($\boldsymbol{\mu}$P) provides a theoretical safeguard for width-invariant $\Theta(1)$ activation…

Machine Learning · Computer Science 2026-03-06 Tian Xie , Haoming Luo , Haoyu Tang , Yiwen Hu , Jason Klein Liu , Qingnan Ren , Yang Wang , Wayne Xin Zhao , Rui Yan , Bing Su , Chong Luo , Baining Guo

With the growing incorporation of deep neural network (DNN) models into modern software systems, the prohibitive construction costs have become a significant challenge. Model reuse has been widely applied to reduce training costs, but…

Machine Learning · Computer Science 2025-08-18 Xiaohan Bi , Binhang Qi , Hailong Sun , Xiang Gao , Yue Yu , Xiaojun Liang

Optimization with matrix gradient orthogonalization has recently demonstrated impressive results in the training of deep neural networks (Jordan et al., 2024; Liu et al., 2025). In this paper, we provide a theoretical analysis of this…

Machine Learning · Computer Science 2025-04-09 Dmitry Kovalev

Structured pruning and knowledge distillation (KD) are typical techniques for compressing large language models, but it remains unclear how they should be applied at pretraining scale, especially to recent mixture-of-experts (MoE) models.…

Machine Learning · Computer Science 2026-05-19 Shengkun Tang , Zekun Wang , Bo Zheng , Liangyu Wang , Rui Men , Siqi Zhang , Xiulong Yuan , Zihan Qiu , Zhiqiang Shen , Dayiheng Liu

Decomposition-based multi-objective evolutionary algorithms (MOEAs) are widely used for solving multi-objective optimisation problems. However, their effectiveness depends on the consistency between the problems Pareto front shape and the…

Neural and Evolutionary Computing · Computer Science 2025-02-25 Xiaofeng Han , Xiaochen Chu , Tao Chao , Ming Yang , Miqing Li

Large Language Models have achieved impressive performance on reasoning-intensive tasks, yet optimizing their reasoning efficiency remains an open challenge. While Test-Time Scaling (TTS) improves reasoning quality, it often leads to…

Computation and Language · Computer Science 2026-05-26 Hang Yan , Fangzhi Xu , Rongman Xu , Yifei Li , Jian Zhang , Haoran Luo , Xiaobao Wu , Luu Anh Tuan , Haiteng Zhao , Qika Lin , Jun Liu

Societal biases are reflected in large pre-trained language models and their fine-tuned versions on downstream tasks. Common in-processing bias mitigation approaches, such as adversarial training and mutual information removal, introduce…

Machine Learning · Computer Science 2023-06-06 Lukas Hauzenberger , Shahed Masoudian , Deepak Kumar , Markus Schedl , Navid Rekabsaz