English
Related papers

Related papers: Kimi K2: Open Agentic Intelligence

200 papers

Existing mobile device control agents often perform poorly when solving complex tasks requiring long-horizon planning and precise operations, typically due to a lack of relevant task experience or unfamiliarity with skill execution. We…

Artificial Intelligence · Computer Science 2026-03-03 Zhe Wu , Donglin Mo , Hongjin Lu , Junliang Xing , Jianheng Liu , Yuheng Jing , Kai Li , Kun Shao , Jianye Hao , Yuanchun Shi

The Mixture of Experts (MoE) architecture has emerged as a powerful paradigm for scaling large language models (LLMs) while maintaining inference efficiency. However, their enormous memory requirements make them prohibitively expensive to…

Machine Learning · Computer Science 2025-06-24 Zichong Li , Chen Liang , Zixuan Zhang , Ilgee Hong , Young Jin Kim , Weizhu Chen , Tuo Zhao

We present Uni-MoE 2.0 from the Lychee family. As a fully open-source omnimodal large model (OLM), it substantially advances Lychee's Uni-MoE series in language-centric multimodal understanding, reasoning, and generating. Based on the dense…

Computation and Language · Computer Science 2025-11-25 Yunxin Li , Xinyu Chen , Shenyuan Jiang , Haoyuan Shi , Zhenyu Liu , Xuanyu Zhang , Nanhao Deng , Zhenran Xu , Yicheng Ma , Meishan Zhang , Baotian Hu , Min Zhang

Multimodal Large Language Models (MLLMs) based agents have demonstrated remarkable potential in autonomous web navigation. However, handling long-horizon tasks remains a critical bottleneck. Prevailing strategies often rely heavily on…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Dawei Yan , Haokui Zhang , Guangda Huzhang , Yang Li , Yibo Wang , Qing-Guo Chen , Zhao Xu , Weihua Luo , Ying Li , Wei Dong , Chunhua Shen

Multilingual end-to-end (E2E) models have shown great promise in expansion of automatic speech recognition (ASR) coverage of the world's languages. They have shown improvement over monolingual systems, and have simplified training and…

Audio and Speech Processing · Electrical Eng. & Systems 2019-09-13 Anjuli Kannan , Arindrima Datta , Tara N. Sainath , Eugene Weinstein , Bhuvana Ramabhadran , Yonghui Wu , Ankur Bapna , Zhifeng Chen , Seungji Lee

We present Laguna M.1 and Laguna XS.2, two Mixture-of-Experts foundation models built for long-horizon, agentic coding: M.1 has $225.8$B total parameters ($23.4$B activated per token) and XS.2 has $33.4$B total ($3$B activated). Both models…

Artificial Intelligence · Computer Science 2026-05-28 Julien Abadji , Marah Abdin , Connor Adams , Eric Alcaide , Mustafa Altun , Michele Artoni , Junze Bao , Uday Barar , Vassilis Bekiaris , Arkadii Bessonov , Benjamin Bütikofer , Jonathan Chang , Yen-Chun Chen , Dmitry Chernenkov , Yang Chi , Filippos Christianos , Fenia Christopoulou , Razvan-Andrei Ciocoiu , Tzachi Cohen , Yohann Coppel , Dmitrii Emelianenko , Brandon Fergerson , Brian Fitzgerald , Matthias Gallé , Alex Golonzovskyi , George Grigorev , Yiyang Hao , Christian Hensel , Jan Huenermann , Ye Ji , Sarthak Joshi , Eiso Kant , Kabir Khandpur , Seonghyeon Kim , Vladimir Kirichenko , Umut Kocasarac , Ilya Kochik , Ivan Komarov , Chaerin Kong , Anurag Koul , François-Joseph Lacroix , Sergei Laktionov , Waren Long , Quentin Malartic , Vadim Markovtsev , Afonso Marques , Robert McHardy , Carlos Mocholí , Dmitry Monakhov , Adam Morris , Martin Muller , Christian Mürtz , Robin Nabel , Thien Nguyen , Rok Novosel , Szymon Ozog , Aalhad Patankar , Aleksei Petrov , Alexandre Piché , Arthur Pignet , Teodor Poncu , Phil Potter , Alexander Rakowski , Pierre-Yves Ritschard , Jay Roberts , Joe Rowell , Piotr Sarna , Pierre-André Savalle , Uladzislau Sazanovich , Nikita Shapovalov , Arsenii Shevchenko , Mikhail Shilkov , Andrei Sokol , Mohamed Soliman , Jack Stephenson , Victor Storchan , Dragos-Constantin Tantaru , Artem Tyurin , Adrian Wälchli , Pengming Wang , Jianxiao Yang , Renat Zayashnikov , Alexander Zelenka Martin , Nikolay Zinov , Caroline Bercier , José Caldeira , Margarida Garcia , Tom George , Kabeer Gharzai , Glenn Hitchcock , Carson Klingenberg , Ivo Pinto , Varun Randery , Noah Smith , Arina Sugako , Jason Warner

Structured pruning and knowledge distillation (KD) are typical techniques for compressing large language models, but it remains unclear how they should be applied at pretraining scale, especially to recent mixture-of-experts (MoE) models.…

Machine Learning · Computer Science 2026-05-19 Shengkun Tang , Zekun Wang , Bo Zheng , Liangyu Wang , Rui Men , Siqi Zhang , Xiulong Yuan , Zihan Qiu , Zhiqiang Shen , Dayiheng Liu

We present Qwen3-Coder-Next, an open-weight language model specialized for coding agents. Qwen3-Coder-Next is an 80-billion-parameter model that activates only 3 billion parameters during inference, enabling strong coding capability with…

Large language models (LLMs) have demonstrated remarkable capabilities across diverse tasks, yet their ability to accurately assess their own confidence remains poorly understood. We present an empirical study investigating whether LLMs…

Computation and Language · Computer Science 2026-03-12 Sudipta Ghosh , Mrityunjoy Panday

We present the technical report for Arcee Trinity Large, a sparse Mixture-of-Experts model with 400B total parameters and 13B activated per token. Additionally, we report on Trinity Nano and Trinity Mini, with Trinity Nano having 6B total…

Mixture-of-Experts (MoE) models, the state-of-the-art in large-scale AI, achieve high quality by sparsely activating parameters. However, their reliance on routing between a few monolithic experts via a top-k mechanism creates a "quality…

Computation and Language · Computer Science 2025-10-23 Xinfeng Xia , Jiacheng Liu , Xiaofeng Hou , Peng Tang , Mingxuan Zhang , Wenfeng Wang , Chao Li

In cross-border e-commerce, search relevance modeling faces the dual challenge of extreme linguistic diversity and fine-grained semantic nuances. Existing approaches typically rely on scaling up a single monolithic Large Language Model…

Information Retrieval · Computer Science 2026-02-04 Ye Liu , Xu Chen , Wuji Chen , Mang Li

In this report, we introduce Qwen2.5, a comprehensive series of large language models (LLMs) designed to meet diverse needs. Compared to previous iterations, Qwen 2.5 has been significantly improved during both the pre-training and…

Operating and maintaining (O&M) large-scale online engine systems (eg, search, recommendation and advertising) demands substantial human effort for release monitoring, alert response, and root cause analysis. Despite the inherent…

Artificial Intelligence · Computer Science 2026-05-12 Bochao Liu , Zhipeng Qian , Yang Zhao , Xinyuan Jiang , Zihan Liang , Yufei Ma , Junpeng Zhuang , Ben Chen , Shuo Yang , Hongen Wan , Yao Wu , Chenyi Lei , Xiao Liang

In this work, we explore a cost-effective framework for multilingual image generation. We find that, unlike models tuned on high-quality images with multilingual annotations, leveraging text encoders pre-trained on widely available, noisy…

Computation and Language · Computer Science 2025-06-06 Sen Xing , Muyan Zhong , Zeqiang Lai , Liangchen Li , Jiawen Liu , Yaohui Wang , Jifeng Dai , Wenhai Wang

The Muon optimizer is consistently faster than Adam in training Large Language Models (LLMs), yet the mechanism underlying its success remains unclear. This paper demystifies this mechanism through the lens of associative memory. By…

Machine Learning · Computer Science 2025-10-07 Shuche Wang , Fengzhuo Zhang , Jiaxiang Li , Cunxiao Du , Chao Du , Tianyu Pang , Zhuoran Yang , Mingyi Hong , Vincent Y. F. Tan

We introduce LongCat-Flash-Prover, a flagship 560-billion-parameter open-source Mixture-of- Experts (MoE) model that advances Native Formal Reasoning in Lean4 through agentic tool-integrated reasoning (TIR). We decompose the native formal…

Neural MMO 2.0 is a massively multi-agent environment for reinforcement learning research. The key feature of this new version is a flexible task system that allows users to define a broad range of objectives and reward signals. We…

Despite the impressive performance of large language models (LLMs) pretrained on vast knowledge corpora, advancing their knowledge manipulation-the ability to effectively recall, reason, and transfer relevant knowledge-remains challenging.…

Computation and Language · Computer Science 2026-01-13 Qitan Lv , Tianyu Liu , Qiaosheng Zhang , Xingcheng Xu , Chaochao Lu

Multimodal large language models (MLLMs) have shown remarkable potential as human-like autonomous language agents to interact with real-world environments, especially for graphical user interface (GUI) automation. However, those GUI agents…

Computation and Language · Computer Science 2024-06-04 Xinbei Ma , Zhuosheng Zhang , Hai Zhao
‹ Prev 1 3 4 5 6 7 10 Next ›