English
Related papers

Related papers: RoboWheel: A Data Engine from Real-World Human Dem…

200 papers

Current digital human studies focusing on lip-syncing and body movement are no longer sufficient to meet the growing industrial demand, while human video generation techniques that support interacting with real-world environments (e.g.,…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Yingying Fan , Quanwei Yang , Kaisiyuan Wang , Hang Zhou , Yingying Li , Haocheng Feng , Errui Ding , Yu Wu , Jingdong Wang

Synthesizing physically plausible articulated human-object interactions (HOI) without 3D/4D supervision remains a fundamental challenge. While recent zero-shot approaches leverage video diffusion models to synthesize human-object…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Zihao Huang , Tianqi Liu , Zhaoxi Chen , Shaocong Xu , Saining Zhang , Lixing Xiao , Zhiguo Cao , Wei Li , Hao Zhao , Ziwei Liu

Monocular 4D human-object interaction (HOI) reconstruction - recovering a moving human and a manipulated object from a single RGB video - remains challenging due to depth ambiguity and frequent occlusions. Existing methods often rely on…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Haoyu Zhang , Wei Zhai , Yuhang Yang , Yang Cao , Zheng-Jun Zha

We present DreamHOI, a novel method for zero-shot synthesis of human-object interactions (HOIs), enabling a 3D human model to realistically interact with any given object based on a textual description. This task is complicated by the…

Computer Vision and Pattern Recognition · Computer Science 2024-09-13 Thomas Hanwen Zhu , Ruining Li , Tomas Jakab

Understanding action correspondence between humans and robots is essential for evaluating alignment in decision-making, particularly in human-robot collaboration and imitation learning within unstructured environments. We propose a…

Robotics · Computer Science 2025-04-17 Azizul Zahid , Jie Fan , Farong Wang , Ashton Dy , Sai Swaminathan , Fei Liu

Human videos offer a scalable way to train robot manipulation policies, but lack the action labels needed by standard imitation learning algorithms. Existing cross-embodiment approaches try to map human motion to robot actions, but often…

We introduce a novel system for human-to-robot trajectory transfer that enables robots to manipulate objects by learning from human demonstration videos. The system consists of four modules. The first module is a data collection module that…

Robotics · Computer Science 2025-10-27 Sai Haneesh Allu , Jishnu Jaykumar P , Ninad Khargonkar , Tyler Summers , Jian Yao , Yu Xiang

Human pose, action, and motion generation are critical for applications in digital humans, character animation, and humanoid robotics. However, many existing methods struggle to produce physically plausible movements that are consistent…

Computer Vision and Pattern Recognition · Computer Science 2025-07-16 Zixi Kang , Xinghan Wang , Yadong Mu

Synthesizing realistic human-object interactions (HOI) in video is challenging due to the complex, instance-specific interaction dynamics of both humans and objects. Incorporating controllability in video generation further adds to the…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Wanyue Zhang , Lin Geng Foo , Thabo Beeler , Rishabh Dabral , Christian Theobalt

Human-Object Interaction (HOI) detection aims to understand the interactions between humans and objects, which plays a curtail role in high-level semantic understanding tasks. However, most works pursue designing better architectures to…

Computer Vision and Pattern Recognition · Computer Science 2023-08-08 Shuman Fang , Shuai Liu , Jie Li , Guannan Jiang , Xianming Lin , Rongrong Ji

Digital human video generation is gaining traction in fields like education and e-commerce, driven by advancements in head-body animation and lip-syncing technologies. However, realistic Hand-Object Interaction (HOI) - the complex dynamics…

Manipulation has long been a challenging task for robots, while humans can effortlessly perform complex interactions with objects, such as hanging a cup on the mug rack. A key reason is the lack of a large and uniform dataset for teaching…

Robotics · Computer Science 2025-06-09 Hongyan Zhi , Peihao Chen , Siyuan Zhou , Yubo Dong , Quanxi Wu , Lei Han , Mingkui Tan

Advances in robotics have been driving the development of human-robot interaction (HRI) technologies. However, accurately perceiving human actions and achieving adaptive control remains a challenge in facilitating seamless coordination…

Robotics · Computer Science 2025-07-10 Mingqi Yuan , Huijiang Wang , Kai-Fung Chu , Fumiya Iida , Bo Li , Wenjun Zeng

World models have become indispensable tools for embodied intelligence, serving as powerful simulators capable of generating realistic robotic videos while addressing critical data scarcity challenges. However, current embodied world models…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Yu Shang , Xin Zhang , Yinzhou Tang , Lei Jin , Chen Gao , Wei Wu , Yong Li

Large-scale real-world robot data collection is a prerequisite for bringing robots into everyday deployment. However, existing pipelines often rely on specialized handheld devices to bridge the embodiment gap, which not only increases…

Robotics · Computer Science 2026-04-10 Yanwen Zou , Chenyang Shi , Wenye Yu , Han Xue , Jun Lv , Ye Pan , Chuan Wen , Cewu Lu

Humanoid whole-body loco-manipulation promises transformative capabilities for daily service and warehouse tasks. While recent advances in general motion tracking (GMT) have enabled humanoids to reproduce diverse human motions, these…

Robotics · Computer Science 2025-10-09 Siheng Zhao , Yanjie Ze , Yue Wang , C. Karen Liu , Pieter Abbeel , Guanya Shi , Rocky Duan

Dexterous manipulation is critical for advancing robot capabilities in real-world applications, yet diverse and high-quality datasets remain scarce. Existing data collection methods either rely on human teleoperation or require significant…

Text-conditioned human motion generation has experienced significant advancements with diffusion models trained on extensive motion capture data and corresponding textual annotations. However, extending such success to 3D dynamic…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Sirui Xu , Ziyin Wang , Yu-Xiong Wang , Liang-Yan Gui
‹ Prev 1 3 4 5 6 7 10 Next ›