English

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Computation and Language 2025-06-17 v1 Machine Learning

Abstract

We introduce MiniMax-M1, the world's first open-weight, large-scale hybrid-attention reasoning model. MiniMax-M1 is powered by a hybrid Mixture-of-Experts (MoE) architecture combined with a lightning attention mechanism. The model is developed based on our previous MiniMax-Text-01 model, which contains a total of 456 billion parameters with 45.9 billion parameters activated per token. The M1 model natively supports a context length of 1 million tokens, 8x the context size of DeepSeek R1. Furthermore, the lightning attention mechanism in MiniMax-M1 enables efficient scaling of test-time compute. These properties make M1 particularly suitable for complex tasks that require processing long inputs and thinking extensively. MiniMax-M1 is trained using large-scale reinforcement learning (RL) on diverse problems including sandbox-based, real-world software engineering environments. In addition to M1's inherent efficiency advantage for RL training, we propose CISPO, a novel RL algorithm to further enhance RL efficiency. CISPO clips importance sampling weights rather than token updates, outperforming other competitive RL variants. Combining hybrid-attention and CISPO enables MiniMax-M1's full RL training on 512 H800 GPUs to complete in only three weeks, with a rental cost of just $534,700. We release two versions of MiniMax-M1 models with 40K and 80K thinking budgets respectively, where the 40K model represents an intermediate phase of the 80K training. Experiments on standard benchmarks show that our models are comparable or superior to strong open-weight models such as the original DeepSeek-R1 and Qwen3-235B, with particular strengths in complex software engineering, tool utilization, and long-context tasks. We publicly release MiniMax-M1 at https://github.com/MiniMax-AI/MiniMax-M1.

Keywords

Cite

@article{arxiv.2506.13585,
  title  = {MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention},
  author = {MiniMax and : and Aili Chen and Aonian Li and Bangwei Gong and Binyang Jiang and Bo Fei and Bo Yang and Boji Shan and Changqing Yu and Chao Wang and Cheng Zhu and Chengjun Xiao and Chengyu Du and Chi Zhang and Chu Qiao and Chunhao Zhang and Chunhui Du and Congchao Guo and Da Chen and Deming Ding and Dianjun Sun and Dong Li and Enwei Jiao and Haigang Zhou and Haimo Zhang and Han Ding and Haohai Sun and Haoyu Feng and Huaiguang Cai and Haichao Zhu and Jian Sun and Jiaqi Zhuang and Jiaren Cai and Jiayuan Song and Jin Zhu and Jingyang Li and Jinhao Tian and Jinli Liu and Junhao Xu and Junjie Yan and Junteng Liu and Junxian He and Kaiyi Feng and Ke Yang and Kecheng Xiao and Le Han and Leyang Wang and Lianfei Yu and Liheng Feng and Lin Li and Lin Zheng and Linge Du and Lingyu Yang and Lunbin Zeng and Minghui Yu and Mingliang Tao and Mingyuan Chi and Mozhi Zhang and Mujie Lin and Nan Hu and Nongyu Di and Peng Gao and Pengfei Li and Pengyu Zhao and Qibing Ren and Qidi Xu and Qile Li and Qin Wang and Rong Tian and Ruitao Leng and Shaoxiang Chen and Shaoyu Chen and Shengmin Shi and Shitong Weng and Shuchang Guan and Shuqi Yu and Sichen Li and Songquan Zhu and Tengfei Li and Tianchi Cai and Tianrun Liang and Weiyu Cheng and Weize Kong and Wenkai Li and Xiancai Chen and Xiangjun Song and Xiao Luo and Xiao Su and Xiaobo Li and Xiaodong Han and Xinzhu Hou and Xuan Lu and Xun Zou and Xuyang Shen and Yan Gong and Yan Ma and Yang Wang and Yiqi Shi and Yiran Zhong and Yonghong Duan and Yongxiang Fu and Yongyi Hu and Yu Gao and Yuanxiang Fan and Yufeng Yang and Yuhao Li and Yulin Hu and Yunan Huang and Yunji Li and Yunzhi Xu and Yuxin Mao and Yuxuan Shi and Yuze Wenren and Zehan Li and Zelin Li and Zhanxu Tian and Zhengmao Zhu and Zhenhua Fan and Zhenzhen Wu and Zhichao Xu and Zhihang Yu and Zhiheng Lyu and Zhuo Jiang and Zibo Gao and Zijia Wu and Zijian Song and Zijun Sun},
  journal= {arXiv preprint arXiv:2506.13585},
  year   = {2025}
}

Comments

A technical report from MiniMax. The authors are listed in alphabetical order. We open-source our MiniMax-M1 at https://github.com/MiniMax-AI/MiniMax-M1

R2 v1 2026-07-01T03:19:53.144Z