We present Mamoda2.5, a unified AR-Diffusion framework that seamlessly integrates multimodal understanding and generation within a single architecture. To efficiently enhance the model's generation capability, we equip the Diffusion Transformer backbone with a fine-grained Mixture-of-Experts (MoE) design (128 experts, Top-8 routing), yielding a 25B-parameter model that activates only 3B parameters, significantly reducing training costs while scaling up the model capacity. Mamoda2.5 achieves top-tier generation performance on VBench 2.0 and sets a new record in video editing quality, surpassing evaluated open-source models and matching the performance of current top-tier proprietary models, including the Kling O1 on OpenVE-Bench. Furthermore, we introduce a joint few-step distillation and reinforcement learning framework that compresses the 30-step editing model into a 4-step model and greatly accelerates model inference. Compared to open-source baselines, Mamoda2.5 achieves up to 95.9× faster video editing inference. In real-world applications, Mamoda2.5 has been successfully deployed for content moderation and creative restoration tasks in advertising scenarios, achieving a 98% success rate in internal advertising video editing scenario.
Cite
@article{arxiv.2605.02641,
title = {Mamoda2.5: Enhancing Unified Multimodal Model with DiT-MoE},
author = {Yangming Shi and Shixiang Zhu and Tao Shen and Zhimiao Yu and Dengsheng Chen and Taicai Chen and Yunfei Yang and Juan Zhou and Chen Cheng and Liang Ma and Xibin Wu and Benxuan Yan and Ge Li and Tuoyu Zhang and Dan Li and Chang Liu and Zhenbang Sun},
journal= {arXiv preprint arXiv:2605.02641},
year = {2026}
}