English

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts

Machine Learning 2025-05-27 v1 Artificial Intelligence Computation and Language

Abstract

To tackle the huge computational demand of large foundation models, activation-aware compression techniques without retraining have been introduced. However, since these rely on calibration data, domain shift may arise for unknown downstream tasks. With a computationally efficient calibration, activation-aware pruning can be executed for every prompt adaptively, yet achieving reduced complexity at inference. We formulate it as a mixture of micro-experts, called μ\mu-MoE. Several experiments demonstrate that μ\mu-MoE can dynamically adapt to task/prompt-dependent structured sparsity on the fly.

Keywords

Cite

@article{arxiv.2505.18451,
  title  = {$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts},
  author = {Toshiaki Koike-Akino and Jing Liu and Ye Wang},
  journal= {arXiv preprint arXiv:2505.18451},
  year   = {2025}
}

Comments

10 pages, 4 figures

R2 v1 2026-07-01T02:35:12.540Z