English
Related papers

Related papers: MiniMax-M1: Scaling Test-Time Compute Efficiently …

200 papers

We present MiMo-7B, a large language model born for reasoning tasks, with optimization across both pre-training and post-training stages. During pre-training, we enhance the data preprocessing pipeline and employ a three-stage data mixing…

Mixture of Experts (MoE) models have emerged as a promising paradigm for scaling language models efficiently by activating only a subset of parameters for each input token. In this report, we present dots.llm1, a large-scale MoE model that…

This paper explores the system 1 thinking capability of Large Reasoning Models (LRMs), the intuitive ability to respond efficiently with minimal token usage. While existing LRMs rely on long-chain reasoning and excel at complex tasks, their…

Computation and Language · Computer Science 2026-05-04 Wenyuan Zhang , Shuaiyi Nie , Xinghua Zhang , Zefeng Zhang , Tingwen Liu

MLLMs have been successfully applied to multimodal embedding tasks, yet their generative reasoning capabilities remain underutilized. Directly incorporating chain-of-thought reasoning into embedding learning introduces two fundamental…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Yuchi Wang , Haiyang Yu , Weikang Bian , Jiefeng Long , Xiao Liang , Chao Feng , Hongsheng Li

Large Language Models (LLMs) have demonstrated remarkable performance across various natural language tasks, marking significant strides towards general artificial intelligence. While general artificial intelligence is leveraged by…

Computation and Language · Computer Science 2023-10-31 Yizhe Yang , Huashan Sun , Jiawei Li , Runheng Liu , Yinghao Li , Yuhang Liu , Heyan Huang , Yang Gao

Large Reasoning Models (LRMs) have shown promising accuracy improvements on complex problem-solving tasks. While these models have attained high accuracy by leveraging additional computation at test time, they need to generate long…

Computation and Language · Computer Science 2025-12-16 Coleman Hooper , Sebastian Zhao , Luca Manolache , Sehoon Kim , Michael W. Mahoney , Yakun Sophia Shao , Kurt Keutzer , Amir Gholami

The current generation of large language models (LLMs) is typically designed for broad, general-purpose applications, while domain-specific LLMs, especially in vertical fields like medicine, remain relatively scarce. In particular, the…

In recent years, lightweight large language models (LLMs) have garnered significant attention in the robotics field due to their low computational resource requirements and suitability for edge deployment. However, in task planning --…

Robotics · Computer Science 2025-10-27 Weijie Zhou , Manli Tao , Chaoyang Zhao , Honghui Dong , Ming Tang , Jinqiao Wang

Large Language Models (LLMs) can achieve enhanced complex problem-solving through test-time computing scaling, yet this often entails longer contexts and numerous reasoning token costs. In this paper, we propose an efficient test-time…

Computation and Language · Computer Science 2025-04-02 Zhaojian Yu , Yinghao Wu , Yilun Zhao , Arman Cohan , Xiao-Ping Zhang

In recent years, general-purpose large language models (LLMs) such as GPT, Gemini, Claude, and DeepSeek have advanced at an unprecedented pace. Despite these achievements, their application to finance remains challenging, due to fragmented…

Large language model (LLM)-based multi-agent systems have demonstrated remarkable promise for tackling complex tasks by breaking them down into subtasks that are iteratively planned, executed, observed, and refined. Despite their…

Multiagent Systems · Computer Science 2025-07-15 Enhao Zhang , Erkang Zhu , Gagan Bansal , Adam Fourney , Hussein Mozannar , Jack Gerrits

We introduce the MiniMax-M2 series, a family of Mixture-of-Experts language models built around the principle that mini activations can unleash maximum real-world intelligence. The flagship M2 contains 229.9B total parameters with only 9.8B…

Artificial Intelligence · Computer Science 2026-05-27 MiniMax , : , Aili Chen , Aonian Li , Baichuan Zhou , Bangwei Gong , Binyang Jiang , Boji Dan , Changqing Yu , Chao Wang , Cheng Ma , Cheng Zhong , Cheng Zhu , Chengjun Xiao , Chengyi Yang , Chengyu Du , Chenyang Zhang , Chi Zhang , Chuangyi Huang , Chunhao Zhang , Chunhui Du , Chunyu Zhao , Congchao Guo , Da Chen , Deming Ding , Dianjun Sun , Dongyu Zhang , Enhui Yang , Fei Yu , Guang Zheng , Guodong Zheng , Guohong Li , Haichao Zhu , Haigang Zhou , Haimo Zhang , Han Ding , Hao Zhang , Haohai Sun , Haolin Lyu , Haonan Lu , Haoyu Wang , Huajie Shi , Huiyang Li , Jiacheng Chen , Jian Zhang , Jiaqi Zhuang , Jiaren Cai , Jiaxin Pan , Jiayao Li , Jiayuan Song , Jichuan Zhang , Jie Wang , Jihao Gu , Jin Zhu , Jingwei Dong , Jingyang Li , Jingyu Zhang , Jingze Zhuang , Jinhao Tian , Jinli Liu , Jinyi Hu , Jun Tao , Jun Zhang , Junbin Ruan , Junhao Xu , Junjie Yan , Junteng Liu , Junxian He , Kang Xu , Ke Ji , Ke Yang , Kecheng Xiao , Keyu Duan , Keyu Li , Le Han , Letian Ruan , Li Yuan , Lianfei Yu , Liheng Feng , Lijie Mo , Lin Li , Lingye Bao , Lingyu Yang , Lingyuan Zhou , Loki , Lu Chen , Lunbin Ceng , Ming Li , Ming Zhong , Mingliang Tao , Mingyuan Chi , Mujie Lin , Nan Hu , Ningxin Chen , Peiyin Zhu , Peng Gao , Pengcheng Gao , Pengfei Li , Penglin Li , Pengyu Zhao , Qibin Ren , Qidi Xu , Qihan Ren , Qile Li , Qin Wang , Quanliang Chen , Qunhong Ceng , Rong Tian , Rui Dong , Ruitao Leng , Ruize Zhang , Shanqi Liu , Shaoyu Chen , Sheng Jia , Shun Yao , Shuoran Zhao , Shuqi Yu , Sichen Li , Sicheng Pan , Songquan Zhu , Tengfei Li , Tian Xie , Tiancheng Qin , Tianrun Liang , Wei Liu , Weiqi Xu , Weitao Li , Weixiang Chen , Weiyu Cheng , Weiyu Zhang , Wenhu Chen , Wenqian Zhao , Xiancai Chen , Xiangjun Song , Xiangyuan Wang , Xiao Luo , Xiao Su , Xiaobo Li , Xiaodong Han , Xiaojie Wu , Xihao Song , Xingyi Han , Xinyu Guan , Xuan Lu , Xun Zou , Xunhao Lai , Xutong Li , Yan Gong , Yang Wang , Yang Xu , Yangsen Wang , Ye Tang , Yicheng Chen , Yinran Qiu , Yiqi Shi , Yiting Guo , Yiwen Huang , Yixuan Wang , Yongyi Hu , Yu Gao , Yu Zhang , Yuanxiang Ying , Yuanzhen Zhang , Yubo Wang , Yuchen Song , Yufeng Yang , Yuhang Meng , Yuhang Miao , Yuhao Li , Yujie Liu , Yulin Hu , Yunan Huang , Yunji Li , Yunyi Huang , Yusen Zhang , Yusu Hong , Yutao Xie , Yutong Zhang , Yuwen Liao , Yuxuan Shi , Yuze Wenren , Zebin Li , Zehan Li , Zejian Luo , Zeyu Jin , Zeyuan Sun , Zhanpeng Zhou , Zhaochen Su , Zhendong Li , Zhengmao Zhu , Zhengyuan Peng , Zhenhua Fan , Zhi Zhang , Zhichao Xu , Zhiheng Lv , Zhikang Xu , Zhitao He , Zhiwei He , Zhongyuan Li , Zibo Gao , Zijia Wu , Zijian Song , Zijian Zhou , Zijun Sun , Zishan Huang , Ziying Chen , Ziyue Ge

We introduce Skywork R1V, a multimodal reasoning model extending the an R1-series Large language models (LLM) to visual modalities via an efficient multimodal transfer method. Leveraging a lightweight visual projector, Skywork R1V…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Yi Peng , Peiyu Wang , Xiaokun Wang , Yichen Wei , Jiangbo Pei , Weijie Qiu , Ai Jian , Yunzhuo Hao , Jiachun Pan , Tianyidan Xie , Li Ge , Rongxian Zhuang , Xuchen Song , Yang Liu , Yahui Zhou

We present AM-Thinking-v1, a 32B dense language model that advances the frontier of reasoning, embodying the collaborative spirit of open-source innovation. Outperforming DeepSeek-R1 and rivaling leading Mixture-of-Experts (MoE) models like…

Computation and Language · Computer Science 2025-05-27 Yunjie Ji , Xiaoyu Tian , Sitong Zhao , Haotian Wang , Shuaiting Chen , Yiping Peng , Han Zhao , Xiangang Li

The adoption of long context windows has become a standard feature in Large Language Models (LLMs), as extended contexts significantly enhance their capacity for complex reasoning and broaden their applicability across diverse scenarios.…

Computation and Language · Computer Science 2026-05-20 Wenxuan Li , Chengruidong Zhang , Huiqiang Jiang , Yucheng Li , Yuqing Yang , Lili Qiu

Modern language agents must operate over long-horizon, multi-turn interactions, where they retrieve external information, adapt to observations, and answer interdependent queries. Yet, most LLM systems rely on full-context prompting,…

Computation and Language · Computer Science 2025-07-18 Zijian Zhou , Ao Qu , Zhaoxuan Wu , Sunghwan Kim , Alok Prakash , Daniela Rus , Jinhua Zhao , Bryan Kian Hsiang Low , Paul Pu Liang

Large language models (LLMs) encounter significant adaptation challenges in diverse multitask finetuning. Mixture-of-experts (MoE) provides a promising solution with a dynamic architecture, enabling effective task decoupling. However,…

Machine Learning · Computer Science 2025-05-28 Rongyu Zhang , Yijiang Liu , Huanrui Yang , Shenli Zheng , Dan Wang , Yuan Du , Li Du , Shanghang Zhang

Recent advancements in Multimodal Large Language Models (MLLMs) have demonstrated satisfactory performance across various vision-language tasks. Current approaches for vision and language interaction fall into two categories:…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Feipeng Ma , Yizhou Zhou , Zheyu Zhang , Shilin Yan , Hebei Li , Zilong He , Siying Wu , Fengyun Rao , Yueyi Zhang , Xiaoyan Sun

Large Language Models (LLMs) have demonstrated remarkable capabilities across various applications, but their performance on long-context tasks is often limited by the computational complexity of attention mechanisms. We introduce a novel…

Machine Learning · Computer Science 2025-02-25 Bo Chen , Yingyu Liang , Zhizhou Sha , Zhenmei Shi , Zhao Song

DeepSeek-R1 is a cutting-edge open-source large language model (LLM) developed by DeepSeek, showcasing advanced reasoning capabilities through a hybrid architecture that integrates mixture of experts (MoE), chain of thought (CoT) reasoning,…

Computation and Language · Computer Science 2025-06-03 Jiancheng Ye , Sophie Bronstein , Jiarui Hai , Malak Abu Hashish