English
Related papers

Related papers: MuLoCo: Muon is a practical inner optimizer for Di…

200 papers

Multimodal large language models (MLLMs) have extended the success of large language models (LLMs) to multiple data types, such as image, text and audio, achieving significant performance in various domains, including multimodal…

Computation and Language · Computer Science 2025-06-03 Weiqi Feng , Yangrui Chen , Shaoyu Wang , Yanghua Peng , Haibin Lin , Minlan Yu

We present MiMo-7B, a large language model born for reasoning tasks, with optimization across both pre-training and post-training stages. During pre-training, we enhance the data preprocessing pipeline and employ a three-stage data mixing…

Large Language Models (LLMs) have seen great advance in both academia and industry, and their popularity results in numerous open-source frameworks and techniques in accelerating LLM pre-training, fine-tuning, and inference. Training and…

Performance · Computer Science 2023-12-04 Longteng Zhang , Xiang Liu , Zeyu Li , Xinglin Pan , Peijie Dong , Ruibo Fan , Rui Guo , Xin Wang , Qiong Luo , Shaohuai Shi , Xiaowen Chu

While adaptive gradient methods are the workhorse of modern machine learning, sign-based optimization algorithms such as Lion and Muon have recently demonstrated superior empirical performance over AdamW in training large language models…

Machine Learning · Computer Science 2026-05-11 Dingzhi Yu , Hongyi Tao , Yuanyu Wan , Luo Luo , Lijun Zhang

Learned optimizers (LOs) have the potential to significantly reduce the wall-clock training time of neural networks. However, they can struggle to optimize unseen tasks (meta-generalize), especially when training networks wider than those…

Machine Learning · Computer Science 2026-03-20 Benjamin Thérien , Charles-Étienne Joseph , Boris Knyazev , Edouard Oyallon , Irina Rish , Eugene Belilovsky

The Muon optimizer has demonstrated strong empirical performance in pre-training large language models by performing matrix-level gradient (or momentum) orthogonalization in each layer independently. In this work, we propose TEON, a…

Machine Learning · Computer Science 2026-02-03 Ruijie Zhang , Yequan Zhao , Ziyue Liu , Zhengyang Wang , Dongyang Li , Yupeng Su , Sijia Liu , Zheng Zhang

Training large language models (LLMs) relies on adaptive optimizers such as Adam, which introduce extra operations and require significantly more memory to maintain first- and second-order moments than SGD. While recent works such as…

Machine Learning · Computer Science 2026-05-22 Athanasios Glentis , Jiaxiang Li , Andi Han , Mingyi Hong

Training large language models requires optimization algorithms that are not only statistically effective, but also computationally and memory efficient at extreme scale. Although Adam remains the dominant optimizer for large-scale…

Machine Learning · Computer Science 2026-05-12 Aditya Ranganath

Adversarial training (AT) remains one of the most reliable empirical defenses against adversarial attacks. Its robustness critically depends on how the underlying min-max objective is optimized. In practice, Stochastic Gradient Descent…

Machine Learning · Computer Science 2026-05-27 Jun Yan , Weiquan Huang , Jiankai Zuo , Yujian Mo , Xi Fang , Chengliang Wu , Zeming Wei

We introduce Kimi K2, a Mixture-of-Experts (MoE) large language model with 32 billion activated parameters and 1 trillion total parameters. We propose the MuonClip optimizer, which improves upon Muon with a novel QK-clip technique to…

Machine Learning · Computer Science 2026-02-04 Kimi Team , Yifan Bai , Yiping Bao , Y. Charles , Cheng Chen , Guanduo Chen , Haiting Chen , Huarong Chen , Jiahao Chen , Ningxin Chen , Ruijue Chen , Yanru Chen , Yuankun Chen , Yutian Chen , Zhuofu Chen , Jialei Cui , Hao Ding , Mengnan Dong , Angang Du , Chenzhuang Du , Dikang Du , Yulun Du , Yu Fan , Yichen Feng , Kelin Fu , Bofei Gao , Chenxiao Gao , Hongcheng Gao , Peizhong Gao , Tong Gao , Yuyao Ge , Shangyi Geng , Qizheng Gu , Xinran Gu , Longyu Guan , Haiqing Guo , Jianhang Guo , Xiaoru Hao , Tianhong He , Weiran He , Wenyang He , Yunjia He , Chao Hong , Hao Hu , Yangyang Hu , Zhenxing Hu , Weixiao Huang , Zhiqi Huang , Zihao Huang , Tao Jiang , Zhejun Jiang , Xinyi Jin , Yongsheng Kang , Guokun Lai , Cheng Li , Fang Li , Haoyang Li , Ming Li , Wentao Li , Yang Li , Yanhao Li , Yiwei Li , Zhaowei Li , Zheming Li , Hongzhan Lin , Xiaohan Lin , Zongyu Lin , Chengyin Liu , Chenyu Liu , Hongzhang Liu , Jingyuan Liu , Junqi Liu , Liang Liu , Shaowei Liu , T. Y. Liu , Tianwei Liu , Weizhou Liu , Yangyang Liu , Yibo Liu , Yiping Liu , Yue Liu , Zhengying Liu , Enzhe Lu , Haoyu Lu , Lijun Lu , Yashuo Luo , Shengling Ma , Xinyu Ma , Yingwei Ma , Shaoguang Mao , Jie Mei , Xin Men , Yibo Miao , Siyuan Pan , Yebo Peng , Ruoyu Qin , Zeyu Qin , Bowen Qu , Zeyu Shang , Lidong Shi , Shengyuan Shi , Feifan Song , Jianlin Su , Zhengyuan Su , Lin Sui , Xinjie Sun , Flood Sung , Yunpeng Tai , Heyi Tang , Jiawen Tao , Qifeng Teng , Chaoran Tian , Chensi Wang , Dinglu Wang , Feng Wang , Hailong Wang , Haiming Wang , Jianzhou Wang , Jiaxing Wang , Jinhong Wang , Shengjie Wang , Shuyi Wang , Si Wang , Xinyuan Wang , Yao Wang , Yejie Wang , Yiqin Wang , Yuxin Wang , Yuzhi Wang , Zhaoji Wang , Zhengtao Wang , Zhengtao Wang , Zhexu Wang , Chu Wei , Qianqian Wei , Haoning Wu , Wenhao Wu , Xingzhe Wu , Yuxin Wu , Chenjun Xiao , Jin Xie , Xiaotong Xie , Weimin Xiong , Boyu Xu , Jinjing Xu , L. H. Xu , Lin Xu , Suting Xu , Weixin Xu , Xinran Xu , Yangchuan Xu , Ziyao Xu , Jing Xu , Jing Xu , Junjie Yan , Yuzi Yan , Hao Yang , Xiaofei Yang , Yi Yang , Ying Yang , Zhen Yang , Zhilin Yang , Zonghan Yang , Haotian Yao , Xingcheng Yao , Wenjie Ye , Zhuorui Ye , Bohong Yin , Longhui Yu , Enming Yuan , Hongbang Yuan , Mengjie Yuan , Siyu Yuan , Haobing Zhan , Dehao Zhang , Hao Zhang , Wanlu Zhang , Xiaobin Zhang , Yadong Zhang , Yangkun Zhang , Yichi Zhang , Yizhi Zhang , Yongting Zhang , Yu Zhang , Yutao Zhang , Yutong Zhang , Zheng Zhang , Haotian Zhao , Yikai Zhao , Zijia Zhao , Huabin Zheng , Shaojie Zheng , Longguang Zhong , Jianren Zhou , Xinyu Zhou , Zaida Zhou , Jinguo Zhu , Zhen Zhu , Weiyu Zhuang , Xinxing Zu

Recent advancements in Multimodal Large Language Models (MLLMs) underscore the significance of scalable models and data to boost performance, yet this often incurs substantial computational costs. Although the Mixture of Experts (MoE)…

Artificial Intelligence · Computer Science 2024-05-21 Yunxin Li , Shenyuan Jiang , Baotian Hu , Longyue Wang , Wanqi Zhong , Wenhan Luo , Lin Ma , Min Zhang

Large language models (LLMs) have demonstrated remarkable capabilities across a wide range of tasks. Typically, LLMs are first pre-trained on large corpora and subsequently fine-tuned on task-specific datasets. However, during fine-tuning,…

Machine Learning · Computer Science 2025-10-21 Yupeng Chen , Senmiao Wang , Yushun Zhang , Zhihang Lin , Haozhe Zhang , Weijian Sun , Tian Ding , Ruoyu Sun

Large Language Models (LLMs) have revolutionized Natural Language Processing (NLP) but demand massive GPU resources for training. Lowering the threshold for LLMs training would encourage greater participation from researchers, benefiting…

Computation and Language · Computer Science 2024-06-07 Kai Lv , Yuqing Yang , Tengxiao Liu , Qinghui Gao , Qipeng Guo , Xipeng Qiu

Learning to Optimize (L2O) enhances optimization efficiency with integrated neural networks. L2O paradigms achieve great outcomes, e.g., refitting optimizer, generating unseen solutions iteratively or directly. However, conventional L2O…

Machine Learning · Computer Science 2025-03-17 Mingjia Shi , Ruihan Lin , Xuxi Chen , Yuhao Zhou , Zezhen Ding , Pingzhi Li , Tong Wang , Kai Wang , Zhangyang Wang , Jiheng Zhang , Tianlong Chen

Large Language Models (LLMs) demonstrate strong generalization and reasoning abilities, making them well-suited for complex decision-making tasks such as medical consultation (MC). However, existing LLM-based methods often fail to capture…

Computation and Language · Computer Science 2025-10-13 Zhihao Jia , Mingyi Jia , Junwen Duan , Jianxin Wang

Most decentralized optimization algorithms are handcrafted. While endowed with strong theoretical guarantees, these algorithms generally target a broad class of problems, thereby not being adaptive or customized to specific problem…

Optimization and Control · Mathematics 2024-10-03 Yutong He , Qiulin Shang , Xinmeng Huang , Jialin Liu , Kun Yuan

Diffusion large language models (dLLMs) are compelling alternatives to autoregressive (AR) models because their denoising models operate over the entire sequence. The global planning and iterative refinement features of dLLMs are…

Computation and Language · Computer Science 2025-06-27 Shansan Gong , Ruixiang Zhang , Huangjie Zheng , Jiatao Gu , Navdeep Jaitly , Lingpeng Kong , Yizhe Zhang

This study presents a novel training algorithm depending upon the recently proposed Fitness Dependent Optimizer (FDO). The stability of this algorithm has been verified and performance-proofed in both the exploration and exploitation stages…

Neural and Evolutionary Computing · Computer Science 2022-01-04 Dosti Kh. Abbas , Tarik A. Rashid , Karmand H. Abdallaand Nebojsa Bacanin , Abeer Alsadoon

Multimodal Audio-Language Models (ALMs) can understand and reason over both audio and text. Typically, reasoning performance correlates with model size, with the best results achieved by models exceeding 8 billion parameters. However, no…

Sound · Computer Science 2025-03-12 Soham Deshmukh , Satvik Dixit , Rita Singh , Bhiksha Raj

In this note, we inspect the convergence of a new optimizer for pretraining LLMs, namely the Muon optimizer. Such an optimizer is closely related to a specialized steepest descent method where the update direction is the minimizer of the…

Optimization and Control · Mathematics 2025-06-03 Jiaxiang Li , Mingyi Hong
‹ Prev 1 3 4 5 6 7 10 Next ›