Machine Learning · Computer Science
Efficiently Editing Mixture-of-Experts Models with Compressed Experts
Yifei He, Yang Liu, Chen Liang, Hany Hassan Awadalla
2025-09-04
Machine Learning · Computer Science
Exploiting the Experts: Unauthorized Compression in MoE-LLMs
Pinaki Prasad Guha Neogi, Ahmad Mohammadshirazi, Dheeraj Kulshrestha, Rajiv Ramnath
2025-11-26
Computation and Language · Computer Science
Diversifying the Expert Knowledge for Task-Agnostic Pruning in Sparse Mixture-of-Experts
Zeliang Zhang, Xiaodong Liu, Hao Cheng, Chenliang Xu +1
2025-06-10
Machine Learning · Computer Science
Towards Efficient Mixture of Experts: A Holistic Study of Compression Techniques
Shwai He, Daize Dong, Liang Ding, Ang Li
2025-03-18
Machine Learning · Computer Science
A Provably Effective Method for Pruning Experts in Fine-tuned Sparse Mixture-of-Experts
Mohammed Nowaz Rabbani Chowdhury, Meng Wang, Kaoutar El Maghraoui, Naigang Wang +2
2024-05-31
Machine Learning · Computer Science
Task-Specific Expert Pruning for Sparse Mixture-of-Experts
Tianyu Chen, Shaohan Huang, Yuan Xie, Binxing Jiao +4
2022-06-03
Computation and Language · Computer Science
DiEP: Adaptive Mixture-of-Experts Compression through Differentiable Expert Pruning
Sikai Bai, Haoxi Li, Jie Zhang, Zicong Hong +1
2025-09-22
Machine Learning · Computer Science
MoNE: Replacing Redundant Experts with Lightweight Novices for Structured Pruning of MoE
Geng Zhang, Yuxuan Han, Yuxuan Lou, Yiqi Zhang +2
2026-02-24
Artificial Intelligence · Computer Science
ConMoE: Expert-Pool Consolidation via Prototype Reassignment for MoE Compression
Yilun Yao, Jiaming Pan, Elsie Dai, Peizhuang Cong +2
2026-05-29
Computation and Language · Computer Science
Cluster-Driven Expert Pruning for Mixture-of-Experts Large Language Models
Hongcheng Guo, Juntao Yao, Boyang Wang, Junjia Du +4
2025-04-11
Machine Learning · Computer Science
AIMER: Calibration-Free Task-Agnostic MoE Pruning
Zongfang Liu, Shengkun Tang, Yifan Shen, Huan Wang +1
2026-04-14
Computer Vision and Pattern Recognition · Computer Science
Multilinear Mixture of Experts: Scalable Expert Specialization through Factorization
James Oldfield, Markos Georgopoulos, Grigorios G. Chrysos, Christos Tzelepis +4
2024-10-18
Machine Learning · Computer Science
eMoE: Task-aware Memory Efficient Mixture-of-Experts-Based (MoE) Model Inference
Suraiya Tairin, Shohaib Mahmud, Haiying Shen, Anand Iyer
2025-03-11
Machine Learning · Computer Science
$\infty$-MoE: Generalizing Mixture of Experts to Infinite Experts
Shota Takashiro, Takeshi Kojima, Shohei Taniguchi, Yusuke Iwasawa +1
2026-01-27
Machine Learning · Computer Science
AdapMoE: Adaptive Sensitivity-based Expert Gating and Management for Efficient MoE Inference
Shuzhang Zhong, Ling Liang, Yuan Wang, Runsheng Wang +2
2024-08-21
Computation and Language · Computer Science
Domain-Specific Pruning of Large Mixture-of-Experts Models with Few-shot Demonstrations
Zican Dong, Han Peng, Peiyu Liu, Wayne Xin Zhao +3
2025-05-29
Machine Learning · Computer Science
SlimMoE: Structured Compression of Large MoE Models via Expert Slimming and Distillation
Zichong Li, Chen Liang, Zixuan Zhang, Ilgee Hong +3
2025-06-24
Machine Learning · Computer Science
Condense, Don't Just Prune: Enhancing Efficiency and Performance in MoE Layer Pruning
Mingyu Cao, Gen Li, Jie Ji, Jiaqi Zhang +5
2026-04-21
Computation and Language · Computer Science
Mixture of Neuron Experts
Runxi Cheng, Yuchen Guan, Yucheng Ding, Qingguo Hu +5
2025-10-08
Computation and Language · Computer Science
Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
Naibin Gu, Zhenyu Zhang, Yuchen Feng, Yilong Chen +7
2026-05-12
Machine Learning · Computer Science
REAP the Experts: Why Pruning Prevails for One-Shot MoE compression
Mike Lasby, Ivan Lazarevich, Nish Sinnadurai, Sean Lie +2
2026-05-14
Machine Learning · Computer Science
PreMoE: Proactive Inference for Efficient Mixture-of-Experts
Zehua Pei, Ying Zhang, Hui-Ling Zhen, Tao Yuan +5
2026-04-27