English

HoMMI: Learning Whole-Body Mobile Manipulation from Human Demonstrations

Robotics 2026-05-18 v2

Abstract

We present Whole-Body Mobile Manipulation Interface (HoMMI), a data collection and policy learning framework that learns whole-body mobile manipulation directly from robot-free human demonstrations. We augment UMI interfaces with egocentric sensing to capture the global context required for mobile manipulation, enabling portable, robot-free, and scalable data collection. However, naively incorporating egocentric sensing introduces a larger human-to-robot embodiment gap in both observation and action spaces, making policy transfer difficult. We explicitly bridge this gap with a cross-embodiment hand-eye policy design, including an embodiment agnostic visual representation; a relaxed head action representation; and a whole-body controller that realizes hand-eye trajectories through coordinated whole-body motion under robot-specific physical constraints. Together, these enable long-horizon mobile manipulation tasks requiring bimanual and whole-body coordination, navigation, and active perception. Results are best viewed on: https://hommi-robot.github.io

Keywords

Cite

@article{arxiv.2603.03243,
  title  = {HoMMI: Learning Whole-Body Mobile Manipulation from Human Demonstrations},
  author = {Xiaomeng Xu and Jisang Park and Han Zhang and Eric Cousineau and Aditya Bhat and Jose Barreiros and Dian Wang and Jeannette Bohg and Shuran Song},
  journal= {arXiv preprint arXiv:2603.03243},
  year   = {2026}
}
R2 v1 2026-07-01T11:01:38.248Z