English
Related papers

Related papers: Step 3.5 Flash: Open Frontier-Level Intelligence w…

200 papers

Multi-agent systems provide a powerful way to extend large language models (LLMs) by decomposing a complex task into specialized subtasks handled by different agents. However, their performance is often hindered by error propagation,…

Machine Learning · Computer Science 2026-05-14 Zheng Wang , Yuang Liu , Yangkai Ding

Recently, Mixture-of-Experts (MoE) models have gained attention for efficiently scaling large language models. Although these models are extremely large, their sparse activation enables inference to be performed by accessing only a fraction…

Machine Learning · Computer Science 2026-01-27 Byeongju Kim , Jungwan Lee , Donghyeon Han , Hoi-Jun Yoo , Sangyeob Kim

AI agents execute complex multi-step processes, but current evaluation falls short: outcome metrics report success or failure without explaining why, and process-level approaches struggle to connect failure types to their precise locations…

Mixture of experts (MoE) has recently emerged as an effective framework to advance the efficiency and scalability of machine learning models by softly dividing complex tasks among multiple specialized sub-models termed experts. Central to…

Machine Learning · Statistics 2025-03-06 Huy Nguyen , Nhat Ho , Alessandro Rinaldo

The outstanding capabilities of large language models (LLMs) render them a crucial component in various autonomous agent systems. While traditional methods depend on the inherent knowledge of LLMs without fine-tuning, more recent approaches…

Artificial Intelligence · Computer Science 2024-12-10 Zhirui Deng , Zhicheng Dou , Yutao Zhu , Ji-Rong Wen , Ruibin Xiong , Mang Wang , Weipeng Chen

The Mixture of Experts (MoE) selects a few feed-forward networks (FFNs) per token, achieving an effective trade-off between computational cost and performance. In conventional MoE, each expert is treated as entirely independent, and experts…

Machine Learning · Computer Science 2026-01-27 Shota Takashiro , Takeshi Kojima , Shohei Taniguchi , Yusuke Iwasawa , Yutaka Matsuo

Real-world model deployment across multiple domains requires multimodal models to operate under two complementary regimes: (1) multi-task pretraining, tasks are co-available at design time where related tasks could borrow representational…

Machine Learning · Computer Science 2026-05-12 Xing Han , Shravan Chaudhari , Tanvi Ranade , Rama Chellappa , Suchi Saria

We propose Tensor-Trained Low-Rank Adaptation Mixture of Experts (TT-LoRA MoE), a novel computational framework integrating Parameter-Efficient Fine-Tuning (PEFT) with sparse MoE routing to address scalability challenges in large model…

Machine Learning · Computer Science 2026-01-27 Pradip Kunwar , Minh N. Vu , Maanak Gupta , Mahmoud Abdelsalam , Manish Bhattarai

Multimodal large language models (MLLMs) have achieved impressive performance, but high-resolution visual inputs result in long sequences of visual tokens and substantial inference latency. Reducing redundant visual tokens is critical to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Guoyang Xia , Yifeng Ding , Fengfa Li , Lei Ren , Wei Chen , Fangxiang Feng , Xiaojie Wang

MeanFlow (MF) is a diffusion-motivated generative model that enables efficient few-step generation by learning long jumps directly from noise to data. In practice, it is often used as a latent MF by leveraging the pre-trained Stable…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Zheyuan Hu , Chieh-Hsin Lai , Ge Wu , Yuki Mitsufuji , Stefano Ermon

AI agent inference is driving an inference heavy datacenter future and exposes bottlenecks beyond compute - especially memory capacity, memory bandwidth and high-speed interconnect. We introduce two metrics - Operational Intensity (OI) and…

Artificial Intelligence · Computer Science 2026-01-30 Yiren Zhao , Junyi Liu

AI support of collaborative interactions entails mediating potential misalignment between interlocutor beliefs. Common preference alignment methods like DPO excel in static settings, but struggle in dynamic collaborative tasks where the…

Computation and Language · Computer Science 2025-05-27 Abhijnan Nath , Carine Graff , Andrei Bachinin , Nikhil Krishnaswamy

In February this year Google proposed a new Transformer variant called FLASH, which has a faster speed, lower VRAM footprint and better performance. This is achieved by designing a performant layer named GAU (Gated Attention Unit), which…

Computation and Language · Computer Science 2022-05-19 Zhenjie Liu

Mixture-of-Experts (MoE) language models can reduce computational costs by 2-4$\times$ compared to dense models without sacrificing performance, making them more efficient in computation-bounded scenarios. However, MoE models generally…

Machine Learning · Computer Science 2024-04-09 Bowen Pan , Yikang Shen , Haokun Liu , Mayank Mishra , Gaoyuan Zhang , Aude Oliva , Colin Raffel , Rameswar Panda

Recent multimodal image generators such as GPT-4o, Gemini 2.0 Flash, and Gemini 2.5 Pro excel at following complex instructions, editing images and maintaining concept consistency. However, they are still evaluated by disjoint toolkits:…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Hang Hua , Ziyun Zeng , Yizhi Song , Yunlong Tang , Liu He , Daniel Aliaga , Wei Xiong , Jiebo Luo

We introduce Kimi K2.5, an open-source multimodal agentic model designed to advance general agentic intelligence. K2.5 emphasizes the joint optimization of text and vision so that two modalities enhance each other. This includes a series of…

Computation and Language · Computer Science 2026-02-04 Kimi Team , Tongtong Bai , Yifan Bai , Yiping Bao , S. H. Cai , Yuan Cao , Y. Charles , H. S. Che , Cheng Chen , Guanduo Chen , Huarong Chen , Jia Chen , Jiahao Chen , Jianlong Chen , Jun Chen , Kefan Chen , Liang Chen , Ruijue Chen , Xinhao Chen , Yanru Chen , Yanxu Chen , Yicun Chen , Yimin Chen , Yingjiang Chen , Yuankun Chen , Yujie Chen , Yutian Chen , Zhirong Chen , Ziwei Chen , Dazhi Cheng , Minghan Chu , Jialei Cui , Jiaqi Deng , Muxi Diao , Hao Ding , Mengfan Dong , Mengnan Dong , Yuxin Dong , Yuhao Dong , Angang Du , Chenzhuang Du , Dikang Du , Lingxiao Du , Yulun Du , Yu Fan , Shengjun Fang , Qiulin Feng , Yichen Feng , Garimugai Fu , Kelin Fu , Hongcheng Gao , Tong Gao , Yuyao Ge , Shangyi Geng , Chengyang Gong , Xiaochen Gong , Zhuoma Gongque , Qizheng Gu , Xinran Gu , Yicheng Gu , Longyu Guan , Yuanying Guo , Xiaoru Hao , Weiran He , Wenyang He , Yunjia He , Chao Hong , Hao Hu , Jiaxi Hu , Yangyang Hu , Zhenxing Hu , Ke Huang , Ruiyuan Huang , Weixiao Huang , Zhiqi Huang , Tao Jiang , Zhejun Jiang , Xinyi Jin , Yu Jing , Guokun Lai , Aidi Li , C. Li , Cheng Li , Fang Li , Guanghe Li , Guanyu Li , Haitao Li , Haoyang Li , Jia Li , Jingwei Li , Junxiong Li , Lincan Li , Mo Li , Weihong Li , Wentao Li , Xinhang Li , Xinhao Li , Yang Li , Yanhao Li , Yiwei Li , Yuxiao Li , Zhaowei Li , Zheming Li , Weilong Liao , Jiawei Lin , Xiaohan Lin , Zhishan Lin , Zichao Lin , Cheng Liu , Chenyu Liu , Hongzhang Liu , Liang Liu , Shaowei Liu , Shudong Liu , Shuran Liu , Tianwei Liu , Tianyu Liu , Weizhou Liu , Xiangyan Liu , Yangyang Liu , Yanming Liu , Yibo Liu , Yuanxin Liu , Yue Liu , Zhengying Liu , Zhongnuo Liu , Enzhe Lu , Haoyu Lu , Zhiyuan Lu , Junyu Luo , Tongxu Luo , Yashuo Luo , Long Ma , Yingwei Ma , Shaoguang Mao , Yuan Mei , Xin Men , Fanqing Meng , Zhiyong Meng , Yibo Miao , Minqing Ni , Kun Ouyang , Siyuan Pan , Bo Pang , Yuchao Qian , Ruoyu Qin , Zeyu Qin , Jiezhong Qiu , Bowen Qu , Zeyu Shang , Youbo Shao , Tianxiao Shen , Zhennan Shen , Juanfeng Shi , Lidong Shi , Shengyuan Shi , Feifan Song , Pengwei Song , Tianhui Song , Xiaoxi Song , Hongjin Su , Jianlin Su , Zhaochen Su , Lin Sui , Jinsong Sun , Junyao Sun , Tongyu Sun , Flood Sung , Yunpeng Tai , Chuning Tang , Heyi Tang , Xiaojuan Tang , Zhengyang Tang , Jiawen Tao , Shiyuan Teng , Chaoran Tian , Pengfei Tian , Ao Wang , Bowen Wang , Chensi Wang , Chuang Wang , Congcong Wang , Dingkun Wang , Dinglu Wang , Dongliang Wang , Feng Wang , Hailong Wang , Haiming Wang , Hengzhi Wang , Huaqing Wang , Hui Wang , Jiahao Wang , Jinhong Wang , Jiuzheng Wang , Kaixin Wang , Linian Wang , Qibin Wang , Shengjie Wang , Shuyi Wang , Si Wang , Wei Wang , Xiaochen Wang , Xinyuan Wang , Yao Wang , Yejie Wang , Yipu Wang , Yiqin Wang , Yucheng Wang , Yuzhi Wang , Zhaoji Wang , Zhaowei Wang , Zhengtao Wang , Zhexu Wang , Zihan Wang , Zizhe Wang , Chu Wei , Ming Wei , Chuan Wen , Zichen Wen , Chengjie Wu , Haoning Wu , Junyan Wu , Rucong Wu , Wenhao Wu , Yuefeng Wu , Yuhao Wu , Yuxin Wu , Zijian Wu , Chenjun Xiao , Jin Xie , Xiaotong Xie , Yuchong Xie , Yifei Xin , Bowei Xing , Boyu Xu , Jianfan Xu , Jing Xu , Jinjing Xu , L. H. Xu , Lin Xu , Suting Xu , Weixin Xu , Xinbo Xu , Xinran Xu , Yangchuan Xu , Yichang Xu , Yuemeng Xu , Zelai Xu , Ziyao Xu , Junjie Yan , Yuzi Yan , Guangyao Yang , Hao Yang , Junwei Yang , Kai Yang , Ningyuan Yang , Ruihan Yang , Xiaofei Yang , Xinlong Yang , Ying Yang , Yi Yang , Yi Yang , Zhen Yang , Zhilin Yang , Zonghan Yang , Haotian Yao , Dan Ye , Wenjie Ye , Zhuorui Ye , Bohong Yin , Chengzhen Yu , Longhui Yu , Tao Yu , Tianxiang Yu , Enming Yuan , Mengjie Yuan , Xiaokun Yuan , Yang Yue , Weihao Zeng , Dunyuan Zha , Haobing Zhan , Dehao Zhang , Hao Zhang , Jin Zhang , Puqi Zhang , Qiao Zhang , Rui Zhang , Xiaobin Zhang , Y. Zhang , Yadong Zhang , Yangkun Zhang , Yichi Zhang , Yizhi Zhang , Yongting Zhang , Yu Zhang , Yushun Zhang , Yutao Zhang , Yutong Zhang , Zheng Zhang , Chenguang Zhao , Feifan Zhao , Jinxiang Zhao , Shuai Zhao , Xiangyu Zhao , Yikai Zhao , Zijia Zhao , Huabin Zheng , Ruihan Zheng , Shaojie Zheng , Tengyang Zheng , Junfeng Zhong , Longguang Zhong , Weiming Zhong , M. Zhou , Runjie Zhou , Xinyu Zhou , Zaida Zhou , Jinguo Zhu , Liya Zhu , Xinhao Zhu , Yuxuan Zhu , Zhen Zhu , Jingze Zhuang , Weiyu Zhuang , Ying Zou , Xinxing Zu

We study the problem of efficient generative inference for Transformer models, in one of its most challenging settings: large deep models, with tight latency targets and long sequence lengths. Better understanding of the engineering…

AI agents -- powered by reasoning-capable large language models (LLMs) and integrated with tools, data, and web search -- are poised to transform the internet into a \emph{Web of Agents}: a machine-native ecosystem where autonomous agents…

Artificial Intelligence · Computer Science 2025-09-08 Rajesh Tembarai Krishnamachari , Srividya Rajesh

Federated fine-tuning of Mixture-of-Experts (MoE)-based large language models (LLMs) is challenging due to their massive computational requirements and the resource constraints of participants. Existing working attempts to fill this gap…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-10-13 Fahao Chen , Jie Wan , Peng Li , Zhou Su , Dongxiao Yu

The ability for AI agents to "think with images" requires a sophisticated blend of reasoning and perception. However, current open multimodal agents still largely fall short on the reasoning aspect crucial for real-world tasks like…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Kaican Li , Lewei Yao , Jiannan Wu , Tiezheng Yu , Jierun Chen , Haoli Bai , Lu Hou , Lanqing Hong , Wei Zhang , Nevin L. Zhang