MotuBrain: An Advanced World Action Model for Robot Control
Abstract
Vision-Language-Action (VLA) models generalize semantically well but often lack fine-grained modeling of world dynamics. We present MotuBrain, a unified World Action Model that jointly models video and action under a UniDiffuser formulation with a three-stream Mixture-of-Transformers architecture. A single model supports policy learning, world modeling, video generation, inverse dynamics, and joint video-action prediction, while scaling to heterogeneous multimodal data such as video-only, task-agnostic, and cross-embodiment robot data. Building on Motus, MotuBrain further introduces unified multiview modeling, an independent text stream for stronger language-action coupling, a shared cross-embodiment action representation, and an efficient post-training and deployment recipe for long-horizon real-world control. Our inference stack combines step reduction, compilation, FP8 quantization, DiT caching, V2A-style action-only inference, and real-time chunked closed-loop execution, achieving over 50x speedup over a naive baseline and up to 11 Hz inference. Experimentally, MotuBrain achieves 95.8% and 96.1% average success on RoboTwin 2.0 under clean and randomized settings, respectively, attains the strongest reported EWMScore in our WorldArena comparison, and adapts to new humanoid embodiments with only 50--100 trajectories. These results show that unified world action models can scale in generality, predictive accuracy, and real-world deployability.
Cite
@article{arxiv.2604.27792,
title = {MotuBrain: An Advanced World Action Model for Robot Control},
author = {MotuBrain Team and Chendong Xiang and Fan Bao and Haitian Liu and Hengkai Tan and Hongzhe Bi and James Li and Jiabao Liu and Jingrui Pang and Kiro Jing and Louis Liu and Mengchen Cai and Rongxu Cui and Ruowen Zhao and Runqing Wang and Shuhe Huang and Yao Feng and Yinze Rong and Zeyuan Wang and Jun Zhu},
journal= {arXiv preprint arXiv:2604.27792},
year = {2026}
}