English
Related papers

Related papers: Open-world Hand-Object Interaction Video Generatio…

200 papers

Amodal completion, which is the process of inferring the full appearance of objects despite partial occlusions, is crucial for understanding complex human-object interactions (HOI) in computer vision and robotics. Existing methods, such as…

Computer Vision and Pattern Recognition · Computer Science 2025-08-04 Seunggeun Chi , Enna Sachdeva , Pin-Hao Huang , Kwonjoon Lee

Human-Object Interaction (HOI) recognition in videos requires understanding both visual patterns and geometric relationships as they evolve over time. Visual and geometric features offer complementary strengths. Visual features capture…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Tanqiu Qiao , Ruochen Li , Frederick W. B. Li , Yoshiki Kubotani , Shigeo Morishima , Hubert P. H. Shum

This paper focuses on a challenging setting of simultaneously modeling geometry and appearance of hand-object interaction scenes without any object priors. We follow the trend of dynamic 3D Gaussian Splatting based methods, and address…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Hao Tian , Chenyangguang Zhang , Rui Liu , Wen Shen , Xiaolin Qin

Generating interaction-centric videos, such as those depicting humans or robots interacting with objects, is crucial for embodied intelligence, as they provide rich and diverse visual priors for robot learning, manipulation policy training,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Gen Li , Bo Zhao , Jianfei Yang , Laura Sevilla-Lara

Hands are the main medium when people interact with the world. Generating proper 3D motion for hand-object interaction is vital for applications such as virtual reality and robotics. Although grasp tracking or object manipulation synthesis…

Computer Vision and Pattern Recognition · Computer Science 2024-01-30 Yuze Hao , Jianrong Zhang , Tao Zhuo , Fuan Wen , Hehe Fan

Generating talking avatars is a fundamental task in video generation. Although existing methods can generate full-body talking avatars with simple human motion, extending this task to grounded human-object interaction (GHOI) remains an open…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Youliang Zhang , Zhengguang Zhou , Zhentao Yu , Ziyao Huang , Teng Hu , Sen Liang , Guozhen Zhang , Ziqiao Peng , Shunkai Li , Yi Chen , Zixiang Zhou , Yuan Zhou , Qinglin Lu , Xiu Li

Generating realistic 3D Human-Object Interactions (HOI) is a fundamental task for applications ranging from embodied AI to virtual content creation, which requires harmonizing high-level semantic intent with strict low-level physical…

Robotics · Computer Science 2026-05-08 Hao Wang , Shiqi Wang , Qi Liu

Human-object interaction detection (HOID) refers to localizing interactive human-object pairs in images and identifying the interactions. Since there could be an exponential number of object-action combinations, labeled data is limited -…

Computer Vision and Pattern Recognition · Computer Science 2025-07-18 Sandipan Sarma , Agney Talwarr , Arijit Sur

Generating videos in the first-person perspective has broad application prospects in the field of augmented reality and embodied intelligence. In this work, we explore the cross-view video prediction task, where given an exo-centric video,…

Computer Vision and Pattern Recognition · Computer Science 2025-04-17 Jilan Xu , Yifei Huang , Baoqi Pei , Junlin Hou , Qingqiu Li , Guo Chen , Yuejie Zhang , Rui Feng , Weidi Xie

Generating natural hand-object interactions in 3D is challenging as the resulting hand and object motions are expected to be physically plausible and semantically meaningful. Furthermore, generalization to unseen objects is hindered by the…

Computer Vision and Pattern Recognition · Computer Science 2024-12-24 Sammy Christen , Shreyas Hampali , Fadime Sener , Edoardo Remelli , Tomas Hodan , Eric Sauser , Shugao Ma , Bugra Tekin

We propose a novel system for robot-to-human object handover that emulates human coworker interactions. Unlike most existing studies that focus primarily on grasping strategies and motion planning, our system focus on 1. inferring human…

Robotics · Computer Science 2025-03-06 Hanxin Zhang , Abdulqader Dhafer , Zhou Daniel Hao , Hongbiao Dong

Generating photorealistic 3D hand-object interactions (HOIs) from text is important for applications like robotic grasping and AR/VR content creation. In practice, however, achieving both visual fidelity and physical plausibility remains…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Uyoung Jeong , Yihalem Yimolal Tiruneh , Hyung Jin Chang , Seungryul Baek , Kwang In Kim

We propose to forecast future hand-object interactions given an egocentric video. Instead of predicting action labels or pixels, we directly predict the hand motion trajectory and the future contact points on the next active object (i.e.,…

Computer Vision and Pattern Recognition · Computer Science 2022-04-05 Shaowei Liu , Subarna Tripathi , Somdeb Majumdar , Xiaolong Wang

Text-conditioned human motion generation has experienced significant advancements with diffusion models trained on extensive motion capture data and corresponding textual annotations. However, extending such success to 3D dynamic…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Sirui Xu , Ziyin Wang , Yu-Xiong Wang , Liang-Yan Gui

Human-Object Interaction (HOI) detection aims to localize human-object pairs and recognize their interactions in images. Although DETR-based methods have recently emerged as the mainstream framework for HOI detection, they still suffer from…

Computer Vision and Pattern Recognition · Computer Science 2025-08-13 Yuxiao Wang , Wolin Liang , Yu Lei , Weiying Xue , Nan Zhuang , Qi Liu

Recent generative models can synthesize high-quality images, but they often fail to generate humans interacting with objects using their hands. This arises mostly from the model's misunderstanding of such interactions and the hardships of…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Patrick Kwon , Chen Chen , Hanbyul Joo

Modeling 3D human-object interaction (HOI) is a problem of great interest for computer vision and a key enabler for virtual and mixed-reality applications. Existing methods work in a one-way direction: some recover plausible human…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Ilya A. Petrov , Riccardo Marin , Julian Chibane , Gerard Pons-Moll

Hand-object interaction (HOI) reconstruction and synthesis are becoming central to embodied AI and AR/VR. Yet, despite rapid progress, existing HOI generation research remains fragmented across three disjoint tracks: (1) pose-only synthesis…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Mingju Gao , Kaisen Yang , Huan-ang Gao , Bohan Li , Ao Ding , Wenyi Li , Yangcheng Yu , Jinkun Liu , Shaocong Xu , Yike Niu , Haohan Chi , Hao Chen , Hao Tang , Yu Zhang , Li Yi , Hao Zhao

Recent advances in video generative models enable the synthesis of realistic human-object interaction videos across a wide range of scenarios and object categories, including complex dexterous manipulations that are difficult to capture…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Hyeonwoo Kim , Jeonghwan Kim , Kyungwon Cho , Hanbyul Joo

We present visual action prompts, a unified action representation for action-to-video generation of complex high-DoF interactions while maintaining transferable visual dynamics across domains. Action-driven video generation faces a…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Yuang Wang , Chao Wen , Haoyu Guo , Sida Peng , Minghan Qin , Hujun Bao , Xiaowei Zhou , Ruizhen Hu