English
Related papers

Related papers: Kimi K2: Open Agentic Intelligence

200 papers

Combining existing pre-trained expert LLMs is a promising avenue for scalably tackling large-scale and diverse tasks. However, selecting task-level experts is often too coarse-grained, as heterogeneous tasks may require different expertise…

Computation and Language · Computer Science 2025-07-22 Justin Chih-Yao Chen , Sukwon Yun , Elias Stengel-Eskin , Tianlong Chen , Mohit Bansal

We introduce Ling 2.0, a series reasoning-oriented language foundation built upon the principle that every activation boosts reasoning capability. Designed to scale from tens of billions to one trillion parameters under a unified…

Computation and Language · Computer Science 2025-11-10 Ling Team , Ang Li , Ben Liu , Binbin Hu , Bing Li , Bingwei Zeng , Borui Ye , Caizhi Tang , Changxin Tian , Chao Huang , Chao Zhang , Chen Qian , Chenchen Ju , Chenchen Li , Chengfu Tang , Chilin Fu , Chunshao Ren , Chunwei Wu , Cong Zhang , Cunyin Peng , Dafeng Xu , Daixin Wang , Dalong Zhang , Dingnan Jin , Dingyuan Zhu , Dongke Hu , Fangzheng Zhao , Feifan Wu , Feng Zhu , Gangshan Wang , Haitao Zhang , Hailin Zhao , Hanxiao Zhang , Hanzi Wang , Hao Qian , Haoyi Yu , Heng Zhang , Hongliang Zhang , Hongzhi Luan , Huirong Dong , Huizhong Li , Jia Li , Jia Liu , Jialong Zhu , Jian Sha , Jianping Wei , Jiaolong Yang , Jieyue Ma , Jiewei Wu , Jinjing Huang , Jingyun Tian , Jingyuan Zhang , Jinquan Sun , Juanhui Tu , Jun Liu , Jun Xu , Jun Zhou , Junjie Ou , Junpeng Fang , Kaihong Zhang , Kaiqin Hu , Ke Shi , Kun Tang , Kunlong Chen , Lanyin Mei , Lei Liang , Lei Xu , Libo Zhang , Lin Ju , Lin Yuan , Ling Zhong , Lintao Ma , Lu Liu , Lu Yu , Lun Cai , Meiqi Zhu , Mengying Li , Min Chen , Minghao Xue , Minghong Cai , Mingming Yin , Peijie Jiang , Peilong Zhao , Pingping Liu , Qian Zhao , Qing Cui , Qingxiang Huang , Qingyuan Yang , Quankun Yu , Shaowei Wei , Shijie Lian , Shoujian Zheng , Shun Song , Shungen Zhang , Shuo Zhang , Siyuan Li , Song Liu , Ting Guo , Tong Zhao , Wanli Gu , Weichang Wu , Weiguang Han , Wenjing Fang , Wubin Wang , Xiang Shu , Xiao Shi , Xiaoshun Lan , Xiaolu Zhang , Xiaqing Sun , Xin Zhao , Xingyu Lu , Xiong Xu , Xudong Wang , Xudong Wang , Xuemin Yang , Yajie Yang , Yang Xiang , Yanzhe Li , Yi Zhang , Yilong Wang , Yingxue Li , Yongzhen Guo , Yuzhuo Fu , Yuanyuan Wang , Yue Yang , Yue Yu , Yufeng Deng , Yun Zhang , Yunfei Yu , Yuqi Zhang , Yuxiao He , Zengke Gui , Zhaoxin Huan , Zhaoyang Wang , Zhibo Zhu , Zhihao Wang , Zhiqiang Zhang , Zhoufei Wang , Zihang Zeng , Ziqi Liu , Zitao Xuan , Zuoli Tang

Large Language Models (LLMs) based autonomous agents demonstrate multifaceted capabilities to contribute substantially to economic production. However, existing benchmarks remain focused on single agentic capability, failing to capture…

Artificial Intelligence · Computer Science 2026-04-24 Keyu Li , Junhao Shi , Yang Xiao , Mohan Jiang , Jie Sun , Yunze Wu , Dayuan Fu , Shijie Xia , Xiaojie Cai , Tianze Xu , Weiye Si , Wenjie Li , Dequan Wang , Pengfei Liu

Efficient materials discovery requires reducing costly first-principles calculations for training machine-learned interatomic potentials (MLIPs). We develop an active learning (AL) framework that iteratively selects informative structures…

Machine Learning · Computer Science 2026-01-22 Mohammed Azeez Khan , Aaron D'Souza , Vijay Choyal

Multi-agent systems (MAS) built on large language models (LLMs) offer a promising path toward solving complex, real-world tasks that single-agent systems often struggle to manage. While recent advancements in test-time scaling (TTS) have…

Artificial Intelligence · Computer Science 2025-08-20 Can Jin , Hongwu Peng , Qixin Zhang , Yujin Tang , Dimitris N. Metaxas , Tong Che

We introduce Yuan3.0 Flash, an open-source Mixture-of-Experts (MoE) MultiModal Large Language Model featuring 3.7B activated parameters and 40B total parameters, specifically designed to enhance performance on enterprise-oriented tasks…

While Multimodal Large Language Models (MLLMs) demonstrate remarkable capabilities on static images, they often fall short in comprehending dynamic, information-dense short-form videos, a dominant medium in today's digital landscape. To…

Large language models (LLMs) have demonstrated remarkable performance on a variety of natural language tasks based on just a few examples of natural language instructions, reducing the need for extensive feature engineering. However, most…

This paper describes the architecture and systems built towards solving the SemEval 2023 Task 2: MultiCoNER II (Multilingual Complex Named Entity Recognition) [1]. We evaluate two approaches (a) a traditional Conditional Random Fields model…

Computation and Language · Computer Science 2024-01-02 Kiran Voderhobli Holla , Chaithanya Kumar , Aryan Singh

While reasoning has become a central capability of large language models (LLMs), the reasoning patterns required for different scenarios are often misaligned. Mathematical reasoning typically relies on intrinsic logic to solve closed-world…

Artificial Intelligence · Computer Science 2026-05-12 Junjian Wang , Xin Zhou , Qiran Xu , Kun Zhan

This research combines Knowledge Distillation (KD) and Mixture of Experts (MoE) to develop modular, efficient multilingual language models. Key objectives include evaluating adaptive versus fixed alpha methods in KD and comparing modular…

Artificial Intelligence · Computer Science 2024-07-30 Mohammed Al-Maamari , Mehdi Ben Amor , Michael Granitzer

Multimodal Mixture-of-Experts (MoE) models offer a promising path toward scalable and efficient large vision-language systems. However, existing approaches rely on rigid routing strategies (typically activating a fixed number of experts per…

Machine Learning · Computer Science 2025-11-25 Yuting Gao , Wang Lan , Hengyuan Zhao , Linjiang Huang , Si Liu , Qingpei Guo

Mixture-of-Experts (MoE) effectively scales large language models (LLMs) and vision-language models (VLMs) by increasing capacity through sparse activation. However, preloading all experts into memory and activating multiple experts per…

Machine Learning · Computer Science 2025-10-14 Wei Huang , Yue Liao , Yukang Chen , Jianhui Liu , Haoru Tan , Si Liu , Shiming Zhang , Shuicheng Yan , Xiaojuan Qi

Developing expressive and responsive conversational digital humans is a cornerstone of next-generation human-computer interaction. While large language models (LLMs) have significantly enhanced dialogue capabilities, most current systems…

Machine Learning · Computer Science 2026-02-06 Xiaolin Hu , Hang Yuan , Xinzhu Sang , Binbin Yan , Zhou Yu , Cong Huang , Kai Chen

Muon has emerged as a promising optimizer for large-scale foundation model pre-training by exploiting the matrix structure of neural network updates through iterative orthogonalization. However, its practical efficiency is limited by the…

Machine Learning · Computer Science 2026-04-14 Ziyue Liu , Ruijie Zhang , Zhengyang Wang , Yequan Zhao , Yupeng Su , Zi Yang , Zheng Zhang

The mixture of experts (MoE) model is a sparse variant of large language models (LLMs), designed to hold a better balance between intelligent capability and computational overhead. Despite its benefits, MoE is still too expensive to deploy…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-04-23 Haodong Wang , Qihua Zhou , Zicong Hong , Song Guo

We introduce MLE-bench, a benchmark for measuring how well AI agents perform at machine learning engineering. To this end, we curate 75 ML engineering-related competitions from Kaggle, creating a diverse set of challenging tasks that test…

Reinforcement learning (RL) agent development traditionally requires substantial expertise and iterative effort, often leading to high failure rates and limited accessibility. This paper introduces Agent$^2$, an LLM-driven…

Artificial Intelligence · Computer Science 2025-10-01 Yuan Wei , Xiaohan Shan , Ran Miao , Jianmin Li

Scale has opened new frontiers in natural language processing, but at a high cost. In response, by learning to only activate a subset of parameters in training and inference, Mixture-of-Experts (MoE) have been proposed as an energy…

Computation and Language · Computer Science 2024-08-09 Xingchen Song , Di Wu , Binbin Zhang , Dinghao Zhou , Zhendong Peng , Bo Dang , Fuping Pan , Chao Yang

Mixture-of-Experts (MoE) models mostly use a router to assign tokens to specific expert modules, activating only partial parameters and often outperforming dense models. We argue that the separation between the router's decision-making and…

Computation and Language · Computer Science 2025-06-02 Ang Lv , Ruobing Xie , Yining Qian , Songhao Wu , Xingwu Sun , Zhanhui Kang , Di Wang , Rui Yan
‹ Prev 1 4 5 6 7 8 10 Next ›