English

MiM-DiT: MoE in MoE with Diffusion Transformers for All-in-One Image Restoration

Computer Vision and Pattern Recognition 2026-03-04 v1

Abstract

All-in-one image restoration is challenging because different degradation types, such as haze, blur, noise, and low-light, impose diverse requirements on restoration strategies, making it difficult for a single model to handle them effectively. In this paper, we propose a unified image restoration framework that integrates a dual-level Mixture-of-Experts (MoE) architecture with a pretrained diffusion model. The framework operates at two levels: the Inter-MoE layer adaptively combines expert groups to handle major degradation types, while the Intra-MoE layer further selects specialized sub-experts to address fine-grained variations within each type. This design enables the model to achieve coarse-grained adaptation across diverse degradation categories while performing fine-grained modulation for specific intra-class variations, ensuring both high specialization in handling complex, real-world corruptions. Extensive experiments demonstrate that the proposed method performs favorably against the state-of-the-art approaches on multiple image restoration task.

Keywords

Cite

@article{arxiv.2603.02710,
  title  = {MiM-DiT: MoE in MoE with Diffusion Transformers for All-in-One Image Restoration},
  author = {Lingshun Kong and Jiawei Zhang and Zhengpeng Duan and Xiaohe Wu and Yueqi Yang and Xiaotao Wang and Dongqing Zou and Lei Lei and Jinshan Pan},
  journal= {arXiv preprint arXiv:2603.02710},
  year   = {2026}
}

Comments

Project website: https://github.com/kkkls/MIM-DiT