Ming-Flash-Omni: A Sparse, Unified Architecture for Multimodal Perception and Generation
Abstract
We propose Ming-Flash-Omni, an upgraded version of Ming-Omni, built upon a sparser Mixture-of-Experts (MoE) variant of Ling-Flash-2.0 with 100 billion total parameters, of which only 6.1 billion are active per token. This architecture enables highly efficient scaling (dramatically improving computational efficiency while significantly expanding model capacity) and empowers stronger unified multimodal intelligence across vision, speech, and language, representing a key step toward Artificial General Intelligence (AGI). Compared to its predecessor, the upgraded version exhibits substantial improvements across multimodal understanding and generation. Notably, it achieves strong performance on vision-language understanding benchmarks, with overall scores on par with Gemini 2.5 Pro, and enables seamless switching among multimodal tasks in multi-turn interactions. In speech, it achieves strong performance in contextual and dialect-aware ASR while enabling joint, continuous-generation of speech, sound, and music. In vision, it introduces generative semantic segmentation that achieves competitive standalone performance and enhances spatial control and editing consistency, alongside marked improvements in identity preservation, and high-fidelity in-image text rendering. Together, these capabilities demonstrate that a single unified model can serve as a practical foundation for general-purpose multimodal intelligence.
Keywords
Cite
@article{arxiv.2510.24821,
title = {Ming-Flash-Omni: A Sparse, Unified Architecture for Multimodal Perception and Generation},
author = {Inclusion AI and : and Bowen Ma and Cheng Zou and ChengKun Du and Canxiang Yan and Chunxiang Jin and Chunjie Shen and Chenyu Lian and Chengxiang Fan and Dandan Zheng and Fudong Wang and Furong Xu and Guangming Yao and Haohao Liu and Han Peng and Jun Zhou and Junluan Xia and Jingdong Chen and Jianing Li and Jianxin Sun and Jianjiang Zhu and Jianping Jiang and Jinpeng Ou and Jun Peng and Jin Peng and Kaixiang Ji and Li Tang and Libin Wang and Lixiang Ru and Longhua Tan and Lu Ma and Lan Wang and Mochen Bai and Minghong Cai and Mingxue Yang and Ning Gao and Qingpei Guo and Qinglong Zhang and Qiang Xu and Qin Zhao and Rui Liu and Ruijie Xiong and Ruobing Zheng and Sirui Gao and Shaoxiong Lin and Tao Zhang and Tianqi Li and Tinghao Liu and Tongli Wang and Taoye Huang and Weilong Chai and Xiaomei Wang and Xiaolong Wang and Xiaojian Liu and Xiao Lu and Xiaoyu Li and Xingning Dong and Xuzheng Yu and Xuezhi Wang and Yi Yuan and Yuting Gao and Yuting Xiao and Yunxiao Sun and Yipeng Chen and Yifan Mao and Yifei Wu and Yongjie Lyu and Yingying Zhang and YuQian Li and Ziping Ma and Zhiqiang Fang and Zhihao Qiu and Ziyuan Huang and Zizheng Yang and Zhengyu He},
journal= {arXiv preprint arXiv:2510.24821},
year = {2026}
}
Comments
18 pages, 5 figures