$ω$-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation
Abstract
Humanoid household tasks often require concurrent loco-manipulation, where the robot must move, adjust posture, maintain balance, and manipulate objects as a single coordinated behavior. Yet existing humanoid policies typically decompose locomotion and manipulation, while recent world-action models remain either arm-centric or video-centered. We present -0, a latent predictive whole-body world-action model for real-world humanoid concurrent loco-manipulation. Given a language instruction, current visual observation, and robot proprioceptive state, -0 directly predicts controller-compatible whole-body action latents for real-robot execution. Rather than reconstructing future videos, -0 learns compact future observation embeddings as a lightweight predictive objective, coupling latent visual foresight with diffusion-based whole-body action generation. The model supports egocentric RGB, exocentric RGB, and exocentric depth inputs, and leverages controller-based simulation replay to ground human/public visual-motion priors into robot-executable action latents. We further collect -HOME, a 40+ hour real-world household humanoid dataset with synchronized multi-view observations, whole-body SMPL motions, robot states, and action latents. Real-world experiments on 11 household tasks demonstrate that a single -0 model can produce smooth manipulate-while-moving behaviors and consistently outperform representative imitation learning, VLA, humanoid, and WAM baselines.
Cite
@article{arxiv.2608.06375,
title = {$ω$-0: A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation},
author = {Zhe Li and Zhenzhe Zhang and Yangyang Wei and Wenjie Zhang and Xichen Yuan and Peiyuan Zhi and Gen Li and Xinying Guo and Fengjie Gao and Jianfei Yang and Shanghang Zhang},
journal= {arXiv preprint arXiv:2608.06375},
year = {2026}
}