Humans achieve complex manipulation through coordinated whole-body control, whereas most Vision-Language-Action (VLA) models treat robot body parts largely independently, making high-DoF humanoid control challenging and often unstable. We present HEX, a state-centric framework for coordinated manipulation on full-sized bipedal humanoid robots. HEX introduces a humanoid-aligned universal state representation for scalable learning across heterogeneous embodiments, and incorporates a Mixture-of-Experts Unified Proprioceptive Predictor to model whole-body coordination and temporal motion dynamics from large-scale multi-embodiment trajectory data. To efficiently capture temporal visual context, HEX uses lightweight history tokens to summarize past observations, avoiding repeated encoding of historical images during inference. It further employs a residual-gated fusion mechanism with a flow-matching action head to adaptively integrate visual-language cues with proprioceptive dynamics for action generation. Experiments on real-world humanoid manipulation tasks show that HEX achieves state-of-the-art performance in task success rate and generalization, particularly in fast-reaction and long-horizon scenarios.
@article{arxiv.2604.07993,
title = {HEX: Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation},
author = {Shuanghao Bai and Meng Li and Xinyuan Lv and Jiawei Wang and Xinhua Wang and Fei Liao and Chengkai Hou and Langzhe Gu and Wanqi Zhou and Kun Wu and Ziluo Ding and Zhiyuan Xu and Lei Sun and Shanghang Zhang and Zhengping Che and Jian Tang and Badong Chen},
journal= {arXiv preprint arXiv:2604.07993},
year = {2026}
}