HumanoidUMI: Bridging Robot-Free Demonstrations and Humanoid Whole-Body Manipulation
Abstract
High-quality demonstration data are essential for humanoid robot skill learning, especially for whole-body behaviors that require coordinated perception, locomotion, and manipulation. Existing data-collection methods largely rely on robot teleoperation, which is constrained by hardware accessibility, operator expertise, and limited efficiency. Inspired by the Universal Manipulation Interface (UMI), we propose HumanoidUMI, a portable and robot-free framework for humanoid whole-body data collection. HumanoidUMI uses lightweight VR devices and UMI-inspired grippers to collect sparse human keypoint trajectories, wrist-view observations, and gripper actions. These demonstrations train a high-level policy to predict future keypoints, which are retargeted to robot-native whole-body references and executed by a whole-body controller. Experiments in five real-world scenarios demonstrate the effectiveness of the proposed framework and validate the collected demonstrations for transferable humanoid whole-body skill learning.
Keywords
Cite
@article{arxiv.2606.27239,
title = {HumanoidUMI: Bridging Robot-Free Demonstrations and Humanoid Whole-Body Manipulation},
author = {Hongwu Wang and Chenhao Yu and Youhao Hu and Jiachen Zhang and Yuanyuan Li and Shaqi Luo},
journal= {arXiv preprint arXiv:2606.27239},
year = {2026}
}
Comments
8 pages, 7 figures