English

Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding

Computer Vision and Pattern Recognition 2025-10-09 v1

Abstract

We introduce Lumina-DiMOO, an open-source foundational model for seamless multi-modal generation and understanding. Lumina-DiMOO sets itself apart from prior unified models by utilizing a fully discrete diffusion modeling to handle inputs and outputs across various modalities. This innovative approach allows Lumina-DiMOO to achieve higher sampling efficiency compared to previous autoregressive (AR) or hybrid AR-Diffusion paradigms and adeptly support a broad spectrum of multi-modal tasks, including text-to-image generation, image-to-image generation (e.g., image editing, subject-driven generation, and image inpainting, etc.), as well as image understanding. Lumina-DiMOO achieves state-of-the-art performance on multiple benchmarks, surpassing existing open-source unified multi-modal models. To foster further advancements in multi-modal and discrete diffusion model research, we release our code and checkpoints to the community. Project Page: https://synbol.github.io/Lumina-DiMOO.

Keywords

Cite

@article{arxiv.2510.06308,
  title  = {Lumina-DiMOO: An Omni Diffusion Large Language Model for Multi-Modal Generation and Understanding},
  author = {Yi Xin and Qi Qin and Siqi Luo and Kaiwen Zhu and Juncheng Yan and Yan Tai and Jiayi Lei and Yuewen Cao and Keqi Wang and Yibin Wang and Jinbin Bai and Qian Yu and Dengyang Jiang and Yuandong Pu and Haoxing Chen and Le Zhuo and Junjun He and Gen Luo and Tianbin Li and Ming Hu and Jin Ye and Shenglong Ye and Bo Zhang and Chang Xu and Wenhai Wang and Hongsheng Li and Guangtao Zhai and Tianfan Xue and Bin Fu and Xiaohong Liu and Yu Qiao and Yihao Liu},
  journal= {arXiv preprint arXiv:2510.06308},
  year   = {2025}
}

Comments

33 pages, 13 figures, 10 tables

R2 v1 2026-07-01T06:22:19.138Z