English
Related papers

Related papers: RoboWheel: A Data Engine from Real-World Human Dem…

200 papers

Recovering 4D human-object interaction (HOI) from monocular video is a key step toward scalable 3D content creation, embodied AI, and simulation-based learning. Recent methods can reconstruct temporally coherent human and object…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Yubo Zhao , Yujin Chai , Yunao Dong , Chengfeng Zhao , Zijiao Zeng , Yuan Liu , Chi-Keung Tang

Enabling robust whole-body humanoid-object interaction (HOI) remains challenging due to motion data scarcity and the contact-rich nature. We present HDMI (HumanoiD iMitation for Interaction), a simple and general framework that learns…

Robotics · Computer Science 2025-09-30 Haoyang Weng , Yitang Li , Nikhil Sobanbabu , Zihan Wang , Zhengyi Luo , Tairan He , Deva Ramanan , Guanya Shi

Generating physically realistic humanoid-object interactions (HOI) is a fundamental challenge in robotics. Existing HOI generation approaches, such as diffusion-based models, often suffer from artifacts such as implausible contacts,…

Robotics · Computer Science 2025-08-21 Yuhang Lin , Yijia Xie , Jiahong Xie , Yuehao Huang , Ruoyu Wang , Jiajun Lv , Yukai Ma , Xingxing Zuo

Recent vision-language-action (VLA) models can generate plausible end-effector motions, yet they often fail in long-horizon, contact-rich tasks because the underlying hand-object interaction (HOI) structure is not explicitly represented. An…

Robotics · Computer Science 2026-02-17 Huajian Zeng , Lingyun Chen , Jiaqi Yang , Yuantai Zhang , Fan Shi , Peidong Liu , Xingxing Zuo

Executing reliable Humanoid-Object Interaction (HOI) tasks for humanoid robots is hindered by the lack of generalized control interfaces and robust closed-loop perception mechanisms. In this work, we introduce Perceptive Root-guided…

Robotics · Computer Science 2026-03-03 Yuhang Lin , Jiyuan Shi , Dewei Wang , Jipeng Kong , Yong Liu , Chenjia Bai , Xuelong Li

The development of embodied AI systems is increasingly constrained by the availability and structure of physical interaction data. Despite recent advances in vision-language-action (VLA) models, current pipelines suffer from high data…

Robotics · Computer Science 2026-03-24 Xinhai Sun , Xiang Shi , Menglin Zou , Wenlong Huang

Interaction is one of the core abilities of humanoid robots. However, most existing frameworks focus on non-interactive whole-body control, which limits their practical applicability. In this work, we develop InterReal, a unified…

Robotics · Computer Science 2026-03-10 Dayang Liang , Yuhang Lin , Xinzhe Liu , Jiyuan Shi , Yunlong Liu , Chenjia Bai

We present Human to Humanoid (H2O), a reinforcement learning (RL) based framework that enables real-time whole-body teleoperation of a full-sized humanoid robot with only an RGB camera. To create a large-scale retargeted motion dataset of…

Robotics · Computer Science 2024-03-08 Tairan He , Zhengyi Luo , Wenli Xiao , Chong Zhang , Kris Kitani , Changliu Liu , Guanya Shi

Human-Object Interaction (HOI) video reenactment with realistic motion remains a frontier in expressive digital human creation. Existing approaches primarily handle simple image-plane motion (e.g., in-plane translations), struggling with…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Jinguang Tong , Jinbo Wu , Kaisiyuan Wang , Zhelun Shen , Xuan Huang , Mochu Xiang , Xuesong Li , Yingying Li , Haocheng Feng , Chen Zhao , Hang Zhou , Wei He , Chuong Nguyen , Jingdong Wang , Hongdong Li

Achieving realistic simulations of humans interacting with a wide range of objects has long been a fundamental goal. Extending physics-based motion imitation to complex human-object interactions (HOIs) is challenging due to intricate…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Sirui Xu , Hung Yu Ling , Yu-Xiong Wang , Liang-Yan Gui

Acquiring large-scale, high-fidelity robot demonstration data remains a critical bottleneck for scaling Vision-Language-Action (VLA) models in dexterous manipulation. We propose a Real-Sim-Real data collection and data editing pipeline that…

Robotics · Computer Science 2026-02-10 Jiacheng Fan , Zhiyue Zhao , Yiqian Zhang , Chao Chen , Peide Wang , Hengdi Zhang , Zhengxue Cheng

Loco-manipulation is a fundamental challenge for humanoid robots to achieve versatile interactions in human environments. Although recent studies have made significant progress in humanoid whole-body control, loco-manipulation remains…

Robotics · Computer Science 2025-10-14 Yuhui Fu , Feiyang Xie , Chaoyi Xu , Jing Xiong , Haoqi Yuan , Zongqing Lu

Generalized robots must learn from diverse, large-scale human-object interactions (HOI) to operate robustly in the real world. Monocular internet videos offer a nearly limitless and readily available source of data, capturing an…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Boran Wen , Ye Lu , Sirui Wang , Keyan Wan , Jiahong Zhou , Junxuan Liang , Xinpeng Liu , Bang Xiao , Ruiyang Liu , Yong-Lu Li

Large-scale pre-training using egocentric human videos has proven effective for robot learning. However, the models pre-trained on such data can be suboptimal for robot learning due to the significant visual gap between human hands and…

Robotics · Computer Science 2026-03-17 Guangrun Li , Yaoxu Lyu , Zhuoyang Liu , Chengkai Hou , Jieyu Zhang , Shanghang Zhang

In 3D hand-object interaction (HOI) tasks, estimating precise joint poses of hands and objects from monocular RGB input remains highly challenging due to the inherent geometric ambiguity of RGB images and the severe mutual occlusions that…

Computer Vision and Pattern Recognition · Computer Science 2025-07-21 Yuechen Xie , Haobo Jiang , Jian Yang , Yigong Zhang , Jin Xie

Our work aims to reconstruct hand-object interactions from a single-view image, which is a fundamental but ill-posed task. Unlike methods that reconstruct from videos, multi-view images, or predefined 3D templates, single-view…

Computer Vision and Pattern Recognition · Computer Science 2025-12-22 Yumeng Liu , Xiaoxiao Long , Zemin Yang , Yuan Liu , Marc Habermann , Christian Theobalt , Yuexin Ma , Wenping Wang

Synthesizing human--object interaction (HOI) videos has broad practical value in e-commerce, digital advertising, and virtual marketing. However, current diffusion models, despite their photorealistic rendering capability, still frequently…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Xiangyang Luo , Xiaozhe Xin , Tao Feng , Xu Guo , Meiguang Jin , Junfeng Ma

Humans achieve complex manipulation through coordinated whole-body control, whereas most Vision-Language-Action (VLA) models treat robot body parts largely independently, making high-DoF humanoid control challenging and often unstable. We…

To serve as a scalable data source for embodied AI, world models should act as true simulators that infer interaction dynamics strictly from user actions, rather than mere conditional video generators relying on privileged future object…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Dayou Li , Lulin Liu , Bangya Liu , Shijie Zhou , Jiu Feng , Ziqi Lu , Minghui Zheng , Chenyu You , Zhiwen Fan

Hand manipulating objects is an important interaction motion in our daily activities. We faithfully reconstruct this motion with a single RGBD camera by a novel deep reinforcement learning method to leverage physics. Firstly, we propose…

Computer Vision and Pattern Recognition · Computer Science 2024-05-07 Haoyu Hu , Xinyu Yi , Zhe Cao , Jun-Hai Yong , Feng Xu
‹ Prev 1 2 3 10 Next ›