English
Related papers

Related papers: MiniMax-M1: Scaling Test-Time Compute Efficiently …

200 papers

We introduce MiniMax-01 series, including MiniMax-Text-01 and MiniMax-VL-01, which are comparable to top-tier models while offering superior capabilities in processing longer contexts. The core lies in lightning attention and its efficient…

The evolution of large language models (LLMs) towards applications with ultra-long contexts faces challenges posed by the high computational and memory costs of the Transformer architecture. While existing sparse and linear attention…

Effective reasoning is crucial to solving complex mathematical problems. Recent large language models (LLMs) have boosted performance by scaling test-time computation through long chain-of-thought reasoning. However, transformer-based…

Machine Learning · Computer Science 2025-09-10 Junxiong Wang , Wen-Ding Li , Daniele Paliotta , Daniel Ritter , Alexander M. Rush , Tri Dao

Linear attention is an efficient attention mechanism that has recently emerged as a promising alternative to conventional softmax attention. With its ability to process tokens in linear computational complexities, linear attention, in…

Computation and Language · Computer Science 2024-01-17 Zhen Qin , Weigao Sun , Dong Li , Xuyang Shen , Weixuan Sun , Yiran Zhong

We present Lightning Attention, the first linear attention implementation that maintains a constant training speed for various sequence lengths under fixed memory consumption. Due to the issue with cumulative summation operations (cumsum),…

Computation and Language · Computer Science 2024-06-21 Zhen Qin , Weigao Sun , Dong Li , Xuyang Shen , Weixuan Sun , Yiran Zhong

We present MiMo-V2-Flash, a Mixture-of-Experts (MoE) model with 309B total parameters and 15B active parameters, designed for fast, strong reasoning and agentic capabilities. MiMo-V2-Flash adopts a hybrid attention architecture that…

Computation and Language · Computer Science 2026-01-09 Core Team , Bangjun Xiao , Bingquan Xia , Bo Yang , Bofei Gao , Bowen Shen , Chen Zhang , Chenhong He , Chiheng Lou , Fuli Luo , Gang Wang , Gang Xie , Hailin Zhang , Hanglong Lv , Hanyu Li , Heyu Chen , Hongshen Xu , Houbin Zhang , Huaqiu Liu , Jiangshan Duo , Jianyu Wei , Jiebao Xiao , Jinhao Dong , Jun Shi , Junhao Hu , Kainan Bao , Kang Zhou , Lei Li , Liang Zhao , Linghao Zhang , Peidian Li , Qianli Chen , Shaohui Liu , Shihua Yu , Shijie Cao , Shimao Chen , Shouqiu Yu , Shuo Liu , Tianling Zhou , Weijiang Su , Weikun Wang , Wenhan Ma , Xiangwei Deng , Bohan Mao , Bowen Ye , Can Cai , Chenghua Wang , Chengxuan Zhu , Chong Ma , Chun Chen , Chunan Li , Dawei Zhu , Deshan Xiao , Dong Zhang , Duo Zhang , Fangyue Liu , Feiyu Yang , Fengyuan Shi , Guoan Wang , Hao Tian , Hao Wu , Heng Qu , Hongfei Yi , Hongxu An , Hongyi Guan , Xing Zhang , Yifan Song , Yihan Yan , Yihao Zhao , Yingchun Lai , Yizhao Gao , Yu Cheng , Yuanyuan Tian , Yudong Wang , Zhen Tang , Zhengju Tang , Zhengtao Wen , Zhichao Song , Zhixian Zheng , Zihan Jiang , Jian Wen , Jiarui Sun , Jiawei Li , Jinlong Xue , Jun Xia , Kai Fang , Menghang Zhu , Nuo Chen , Qian Tu , Qihao Zhang , Qiying Wang , Rang Li , Rui Ma , Shaolei Zhang , Shengfan Wang , Shicheng Li , Shuhao Gu , Shuhuai Ren , Sirui Deng , Tao Guo , Tianyang Lu , Weiji Zhuang , Weikang Zhang , Weimin Xiong , Wenshan Huang , Wenyu Yang , Xin Zhang , Xing Yong , Xu Wang , Xueyang Xie , Yilin Jiang , Yixin Yang , Yongzhe He , Yu Tu , Yuanliang Dong , Yuchen Liu , Yue Ma , Yue Yu , Yuxing Xiang , Zhaojun Huang , Zhenru Lin , Zhipeng Xu , Zhiyang Chen , Zhonghua Deng , Zihan Zhang , Zihao Yue

Although Multimodal Large Language Models (MLLMs) have demonstrated remarkable capabilities across diverse tasks, they encounter challenges in terms of reasoning efficiency, large model size and overthinking. However, existing lightweight…

Artificial Intelligence · Computer Science 2025-11-21 Qixiang Yin , Huanjin Yao , Jianghao Chen , Jiaxing Huang , Zhicheng Zhao , Fei Su

In this technical report, we present the Ring-linear model series, specifically including Ring-mini-linear-2.0 and Ring-flash-linear-2.0. Ring-mini-linear-2.0 comprises 16B parameters and 957M activations, while Ring-flash-linear-2.0…

The computational challenges of Large Language Model (LLM) inference remain a significant barrier to their widespread deployment, especially as prompt lengths continue to increase. Due to the quadratic complexity of the attention…

Long-term memory is a cornerstone of human intelligence. Enabling AI to process lifetime-scale information remains a long-standing pursuit in the field. Due to the constraints of full-attention architectures, the effective context length of…

Computation and Language · Computer Science 2026-04-14 Yu Chen , Runkai Chen , Sheng Yi , Xinda Zhao , Xiaohong Li , Jianjin Zhang , Jun Sun , Chuanrui Hu , Yunyun Han , Lidong Bing , Yafeng Deng , Tianqiao Chen

Recent works on large language models (LLMs) have successfully demonstrated the emergence of reasoning capabilities via reinforcement learning (RL). Although recent efforts leverage group relative policy optimization (GRPO) for MLLMs…

Computation and Language · Computer Science 2025-06-18 Shilin Xu , Yanwei Li , Rui Yang , Tao Zhang , Yueyi Sun , Wei Chow , Linfeng Li , Hang Song , Qi Xu , Yunhai Tong , Xiangtai Li , Hao Fei

We present Ring-1T, the first open-source, state-of-the-art thinking model with a trillion-scale parameter. It features 1 trillion total parameters and activates approximately 50 billion per token. Training such models at a…

Recent advances in large language models (LLMs), such as OpenAI-o1 and DeepSeek-R1, have demonstrated the effectiveness of test-time scaling, where extended reasoning processes substantially enhance model performance. Despite this, current…

Computation and Language · Computer Science 2025-03-26 Xiaoyu Tian , Sitong Zhao , Haotian Wang , Shuaiting Chen , Yunjie Ji , Yiping Peng , Han Zhao , Xiangang Li

We present TransNormerLLM, the first linear attention-based Large Language Model (LLM) that outperforms conventional softmax attention-based models in terms of both accuracy and efficiency. TransNormerLLM evolves from the previous linear…

Computation and Language · Computer Science 2024-01-22 Zhen Qin , Dong Li , Weigao Sun , Weixuan Sun , Xuyang Shen , Xiaodong Han , Yunshen Wei , Baohong Lv , Xiao Luo , Yu Qiao , Yiran Zhong

While current Multimodal Large Language Models (MLLMs) have demonstrated proficiency in reasoning tasks such as mathematics and logic, their capacity for long-chain reflective reasoning, a prerequisite for solving complex real-world…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Xiangyu Zhao , Junming Lin , Tianhao Liang , Yifan Zhou , Wenhao Chai , Yuzhe Gu , Weiyun Wang , Kai Chen , Gen Luo , Wenwei Zhang , Junchi Yan , Hua Yang , Haodong Duan , Xue Yang

Large language models have recently evolved from fluent text generation to advanced reasoning across diverse domains, giving rise to reasoning language models. Among these domains, mathematical reasoning serves as a representative benchmark…

Transformer-based models have emerged as one of the most widely used architectures for natural language processing, natural language generation, and image generation. The size of the state-of-the-art models has increased steadily reaching…

Hardware Architecture · Computer Science 2025-01-15 Rya Sanovar , Srikant Bharadwaj , Renee St. Amant , Victor Rühle , Saravan Rajmohan

We present MM1.5, a new family of multimodal large language models (MLLMs) designed to enhance capabilities in text-rich image understanding, visual referring and grounding, and multi-image reasoning. Building upon the MM1 architecture,…

We introduce LongCat-Flash-Thinking-2601, a 560-billion-parameter open-source Mixture-of-Experts (MoE) reasoning model with superior agentic reasoning capability. LongCat-Flash-Thinking-2601 achieves state-of-the-art performance among…

Artificial Intelligence · Computer Science 2026-02-03 Meituan LongCat Team , Anchun Gui , Bei Li , Bingyang Tao , Bole Zhou , Borun Chen , Chao Zhang , Chao Zhang , Chen Gao , Chen Zhang , Chengcheng Han , Chenhui Yang , Chuyu Zhang , Cong Chen , Cunguang Wang , Daoru Pan , Defei Bu , Dengchang Zhao , Di Xiu , Dishan Liu , Dongyu Ru , Dunwei Tu , Fan Wu , Fengcheng Yuan , Fengcun Li , Gang Xu , Guanyu Wu , Guoyuan Lin , Haibin Wang , Hansi Yang , Hao Yang , Haonan Yan , Haoxiang Ma , Haoxing Wen , Hongyan Hao , Hongyin Tang , Hongyu Zang , Hongzhi Ni , Hui Su , Jiacheng Zhang , Jiahong Zhou , Jiahuan Li , Jiaming Wang , Jian Yang , Jianfei Zhang , Jianhao Xu , Jianing Wang , Jiapeng Zhu , Jiaqi Sun , Jiarong Shi , Jiarui Zhao , Jingang Wang , Jinluan Yang , Jinrui Ding , Jinwei Xiao , Jiyuan He , Juncan Xu , Kefeng Zhang , Keheng Wang , Li Wei , Lianhui Ma , Lin Qiu , Lingbing Kong , Lingchuan Liu , Linsen Guo , Mengshen Zhu , Mengxia Shen , Mingyang Zhu , Peiguang Li , Peng Pei , Peng Zhao , Pengcheng Jia , Pengtao Zhang , Ping Liu , Qi Gu , Qiong Huang , Qiyuan Duan , Quanchi Weng , Rongxiang Weng , Rongzhi Zhang , Rumei Li , Shanglin Lei , Shengnan An , Shijun Dai , Shizhe Wu , Shuaikang Liu , Shuang Zhou , Shuo Wang , Songyuan Zhao , Tao Liang , Tianhao Hu , Tianze Chen , Wei Liu , Wei Shi , Wei Wang , Weifeng Tang , Wenjie Shi , Wenlong Zhu , Wentao Chen , Wentao Shi , Xi Su , Xiandi Ma , Xiangcheng Liu , Xiangyu Xi , Xiangyuan Liu , Xiangzhou Huang , Xiao Liu , Xiaodong Cai , Xiaolong Chen , Xiaowei Shi , Xiaoyu Li , Xin Chen , Xingchen Liu , Xuan Huang , Xuezhi Cao , Xunliang Cai , Yan Chen , Yang Bai , Yang Liu , Yang Yang , Yang Zheng , Yanyu Chen , Yaoming Wang , Yaoming Zhu , Yaorui Shi , Yaqi Huo , Yerui Sun , Yi Zhang , Yi-Kai Zhang , Yifan Lu , Yifan Zhao , Yihao Chen , Yitao Zhai , Yongjing Yin , Yongwei Zhou , Youshao Xiao , Yu Wang , Yu Yang , Yuchen Xie , Yuchen Yu , Yuchuan Dai , Yue Xu , Yueqing Sun , Yufei Zhang , Yuhuai Wei , Yulei Qian , Yunfan Liang , Yunke Zhao , Yuwei Jiang , Yuxin Bian , Yuxin Chen , Yuxin Liu , Zeyang Yu , Zhao Yang , Zhengsheng Huang , Zhengyu Chen , Zhijian Liu , Zhikang Xia , Zhimin Lin , Zhiyuan Yao , Zhuofan Chen , Zhuowen Han , Zijian Zhang , Ziran Li , Ziwen Wang , Ziyuan Zhuang
‹ Prev 1 2 3 10 Next ›