中文
相关论文

相关论文: MuCon: Clipped Muon Updates for LLM Training

200 篇论文

Fully decentralized Muon is difficult because its nonlinear matrix-sign operator does not commute with linear gossip averaging. This makes decentralized Muon a structural design problem: in designing the algorithm, one must distinguish…

最优化与控制 · 数学 2026-04-28 Hengrui Zhang , Boao Kong , Jiahe Geng , Zhengyang Huang

We introduce Post-Optimization Model Edit (POME), a new algorithm that enhances the performance of fine-tuned large language models using only their pretrained and fine-tuned checkpoints, without requiring extra data or further…

机器学习 · 计算机科学 2025-10-09 Yong Liu , Di Fu , Yang Luo , Zirui Zhu , Minhao Cheng , Cho-Jui Hsieh , Yang You

The Muon optimizer is consistently faster than Adam in training Large Language Models (LLMs), yet the mechanism underlying its success remains unclear. This paper demystifies this mechanism through the lens of associative memory. By…

Efficient stochastic optimization typically integrates an update direction that performs well in the deterministic regime with a mechanism adapting to stochastic perturbations. While Adam uses adaptive moment estimates to promote stability,…

机器学习 · 计算机科学 2026-02-23 Minxin Zhang , Yuxuan Liu , Hayden Schaeffer

In this article, we explore the use of various matrix norms for optimizing functions of weight matrices, a crucial problem in training large language models. Moving beyond the spectral norm underlying the Muon update, we leverage duals of…

Fine-tuning Large Language Models (LLMs) is essential for adapting pre-trained models to downstream tasks. Yet traditional first-order optimizers such as Stochastic Gradient Descent (SGD) and Adam incur prohibitive memory and computational…

We propose PRISM, an optimizer that enhances first-order spectral descent methods like Muon with partial second-order information. It constructs an efficient, low-rank quasi-second-order preconditioner via innovation-augmented polar…

机器学习 · 计算机科学 2026-02-04 Yujie Yang

Prompt tuning, like CoOp, has recently shown promising vision recognizing and transfer learning ability on various downstream tasks with the emergence of large pre-trained vision-language models like CLIP. However, we identify that existing…

计算机视觉与模式识别 · 计算机科学 2024-07-16 Yongzhu Miao , Shasha Li , Jintao Tang , Ting Wang

In this paper, we focus on the nonconvex-strongly-concave minimax optimization problem (MCC), where the inner maximization subproblem contains constraints that couple the primal variable of the outer minimization problem. We prove that by…

最优化与控制 · 数学 2024-09-02 Xiaoyin Hu , Kim-Chuan Toh , Shiwei Wang , Nachuan Xiao

Zeroth-order (ZO) optimization provides a gradient-free alternative to first-order (FO) methods by estimating gradients via finite differences of function evaluations, and has recently emerged as a memory-efficient paradigm for fine-tuning…

机器学习 · 计算机科学 2026-02-24 Yicheng Lang , Changsheng Wang , Yihua Zhang , Mingyi Hong , Zheng Zhang , Wotao Yin , Sijia Liu

In this paper we consider the problem of distributed nonlinear optimisation of a separable convex cost function over a graph subject to cone constraints. We show how to generalise, using convex analysis, monotone operator theory and…

分布式、并行与集群计算 · 计算机科学 2024-05-16 Richard Heusdens , Guoqiang Zhang

Vision-language pre-trained models have achieved impressive performance on various downstream tasks. However, their large model sizes hinder their utilization on platforms with limited computational resources. We find that directly using…

计算机视觉与模式识别 · 计算机科学 2024-03-13 Haokun Lin , Haoli Bai , Zhili Liu , Lu Hou , Muyi Sun , Linqi Song , Ying Wei , Zhenan Sun

Different gradient-based methods for optimizing overparameterized models can all achieve zero training error yet converge to distinctly different solutions inducing different generalization properties. We provide the first complete…

机器学习 · 计算机科学 2025-12-08 Chen Fan , Mark Schmidt , Christos Thrampoulidis

The recent empirical success of the Muon optimizer has renewed interest in non-Euclidean optimization, typically justified by similarities with second-order methods, and linear minimization oracle (LMO) theory. In this paper, we challenge…

The current discrepancy of theory and experiment observed recently in muonic hydrogen necessitates a reinvestigation of all corrections to contribute to the Lamb shift in muonic hydrogen muH, muonic deuterium muD, the muonic 3He ion, as…

原子物理 · 物理学 2012-07-03 U. D. Jentschura , B. J. Wundt

Training instabilities such as loss spikes are frequently the result of stochastic gradient noise. Because of rare expressions in language training data, and multiple layer composition, the noise impact is heavy-tailed and survives…

机器学习 · 计算机科学 2026-05-28 Zitao Song , Cedar Site Bai , Zhe Zhang , Brian Bullins , David F. Gleich

Prompt learning has become a dominant paradigm for adapting vision-language models (VLMs) such as CLIP to downstream tasks without modifying pretrained weights. While extending prompts to both vision and text encoders across multiple…

计算机视觉与模式识别 · 计算机科学 2026-02-26 Sajjad Ghiasvand , Haniyeh Ehsani Oskouie , Mahnoosh Alizadeh , Ramtin Pedarsani

Universal machine-learned interatomic potentials (U-MLIPs) have demonstrated broad applicability across diverse atomistic systems but often require fine-tuning to achieve task-specific accuracy. While the number of available U-MLIPs and…

计算物理 · 物理学 2025-08-25 Xiaoqing Liu , Kehan Zeng , Zedong Luo , Yangshuai Wang , Teng Zhao , Zhenli Xu

Matrix decomposition is ubiquitous and has applications in various fields like speech processing, data mining and image processing to name a few. Under matrix decomposition, nonnegative matrix factorization is used to decompose a…

最优化与控制 · 数学 2019-05-14 R. Jyothi , P. Babu , R. Bahl

For MIMO systems, due to the deployment of multiple antennas at both the transmitter and the receiver, the design variables e.g., precoders, equalizers, training sequences, etc. are usually matrices. It is well known that matrix operations…

信息论 · 计算机科学 2014-12-04 Chengwen Xing , Shaodan Ma , Yiqing Zhou