English
Related papers

Related papers: Awakening Dormant Experts:Counterfactual Routing t…

200 papers

The cosine router in Mixture of Experts (MoE) has recently emerged as an attractive alternative to the conventional linear router. Indeed, the cosine router demonstrates favorable performance in image and language tasks and exhibits better…

Machine Learning · Statistics 2025-03-06 Huy Nguyen , Pedram Akbarian , Trang Pham , Trang Nguyen , Shujian Zhang , Nhat Ho

Sparse Mixture-of-Experts (MoE) models scale parameters while fixing active computation per token, but the specialization of individual experts remains opaque. In a companion paper we showed that routing topology is quality-neutral: five…

Artificial Intelligence · Computer Science 2026-04-17 Ivan Ternovtsii , Yurii Bilak

Despite MoE models leading many benchmarks, supervised fine-tuning (SFT) for the MoE architectures remains difficult because its router layers are fragile. Methods such as DenseMixer and ESFT mitigate router collapse with dense mixing or…

Machine Learning · Computer Science 2026-04-28 Haoze He , Xingyuan Ding , Xuan Jiang , Xinkai Zou , Alex Cheng , Yibo Zhao , Juncheng Billy Li , Heather Miller

Mixture-of-Experts (MoE) language models route each token to a small subset of experts, but whether the routes selected by a trained top-$k$ router are good ones is rarely evaluated directly. Holding the model fixed, we compare each…

Machine Learning · Computer Science 2026-05-11 Youngsik Yoon , Siwei Wang , Wei Chen , Jungseul Ok

In parameter-efficient fine-tuning, mixture-of-experts (MoE), which involves specializing functionalities into different experts and sparsely activating them appropriately, has been widely adopted as a promising approach to trade-off…

Machine Learning · Computer Science 2025-08-29 Jinyuan Feng , Chaopeng Wei , Tenghai Qiu , Tianyi Hu , Zhiqiang Pu

Referring Multi-Object Tracking (RMOT) faces a fundamental structural contradiction between the high-discriminability demand and the sparse semantic supervision. This mismatch is particularly acute in highly homogeneous scenarios that…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Shukun Jia , Shiyu Hu , Yipei Wang , Ximeng Cheng , Yichao Cao , Xiaobo Lu

We propose Chain-of-Experts (CoE), a new Mixture-of-Experts (MoE) architecture that introduces sequential expert communication within each layer. Unlike traditional MoE models, where experts operate independently in parallel, CoE processes…

Machine Learning · Computer Science 2025-06-25 Zihan Wang , Rui Pan , Jiarui Yao , Robert Csordas , Linjie Li , Lu Yin , Jiajun Wu , Tong Zhang , Manling Li , Shiwei Liu

The demonstrated success of sparsely-gated Mixture-of-Experts (MoE) architectures, exemplified by models such as DeepSeek and Grok, has motivated researchers to investigate their adaptation to diverse domains. In real-world image…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Xiao He , Zhijun Tu , Kun Cheng , Mingrui Zhu , Jie Hu , Nannan Wang , Xinbo Gao

Mixture-of-Experts (MoE) models improve efficiency through sparse activation, but their learned gating functions provide limited insight into routing decisions. This work introduces the Semantic Resonance Architecture (SRA), which routes…

Computation and Language · Computer Science 2026-04-17 Ivan Ternovtsii , Yurii Bilak

Sparse Mixture-of-Experts (MoE) models scale capacity by routing each token to a small subset of experts. However, their routers exhibit a fundamental trade-off: strong load balancing can suppress expert specialization, while aggressive…

Machine Learning · Computer Science 2026-05-12 Gleb Molodtsov , Alexander Miasnikov , Aleksandr Beznosikov

Real-world wireless data are expensive to collect and often lack sufficient expert demonstrations, causing existing offline RL methods to overfit suboptimal behaviors and exhibit unstable performance. To address this issue, we propose CORE,…

Networking and Internet Architecture · Computer Science 2025-12-23 Lipeng Zu , Hansong Zhou , Yu Qian , Shayok Chakraborty , Yukun Yuan , Linke Guo , Xiaonan Zhang

Mixture-of-Experts (MoE) architectures scale large language models efficiently by employing a parametric ``router'' to dispatch tokens to a sparse subset of experts. Typically, this router is trained once and then frozen, rendering routing…

Computation and Language · Computer Science 2026-05-26 Boxuan Lyu , Soichiro Murakami , Hidetaka Kamigaito , Peinan Zhang

Visual Question Answering (VQA) requires models to identify the correct answer options based on both visual and textual evidence. Recent Mixture-of-Experts (MoE) methods improve option reasoning by grouping similar concepts or routing based…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Xiyin Zeng , Yi Lu , Hao Wang

Sparse Mixture of Experts (MoE) models offer a scalable and efficient architecture for training large neural networks by activating only a subset of parameters ("experts") for each input. A learned router computes a distribution over these…

Machine Learning · Computer Science 2025-10-14 Nabil Omi , Siddhartha Sen , Ali Farhadi

Sparse Mixture of Experts (SMoE) enables efficient training of large language models by routing input tokens to a select number of experts. However, training SMoE remains challenging due to the issue of representation collapse. Recent…

Computation and Language · Computer Science 2025-04-01 Giang Do , Hung Le , Truyen Tran

Modern LLMs continue to exhibit significant variance in behavior across languages, such as being able to recall factual information in some languages but not others. While typically studied as a problem to be mitigated, in this work, we…

Computation and Language · Computer Science 2026-03-19 Lucas Bandarkar , Alan Ansell , Trevor Cohn

Recent progress in deep learning has been driven by increasingly large-scale models, but the resulting computational cost has become a critical bottleneck. Sparse Mixture of Experts (MoE) offers an effective solution by activating only a…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Masahiro Kada , Ryota Yoshihashi , Satoshi Ikehata , Rei Kawakami , Ikuro Sato

Large Language Models have demonstrated remarkable capabilities across diverse tasks, yet they frequently generate hallucinations outputs that are fluent but factually incorrect or unsupported. We propose Counterfactual Probing, a novel…

Computation and Language · Computer Science 2025-08-05 Yijun Feng

Mixture of Experts (MoE) pretraining is more scalable than dense Transformer pretraining, because MoEs learn to route inputs to a sparse set of their feedforward parameters. However, this means that MoEs only receive a sparse backward…

Machine Learning · Computer Science 2025-11-05 Ashwinee Panda , Vatsal Baherwani , Zain Sarwar , Benjamin Therien , Sambit Sahu , Tom Goldstein , Supriyo Chakraborty

Sparse mixture of experts provides larger model capacity while requiring a constant computational overhead. It employs the routing mechanism to distribute input tokens to the best-matched experts according to their hidden representations.…

Computation and Language · Computer Science 2022-10-13 Zewen Chi , Li Dong , Shaohan Huang , Damai Dai , Shuming Ma , Barun Patra , Saksham Singhal , Payal Bajaj , Xia Song , Xian-Ling Mao , Heyan Huang , Furu Wei
‹ Prev 1 2 3 10 Next ›