English
Related papers

Related papers: THOM: Generating Physically Plausible Hand-Object …

200 papers

Synthesizing semantic-aware, long-horizon, human-object interaction is critical to simulate realistic human behaviors. In this work, we address the challenging problem of generating synchronized object motion and human motion guided by…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Jiaman Li , Alexander Clegg , Roozbeh Mottaghi , Jiajun Wu , Xavier Puig , C. Karen Liu

Hand-object interaction (HOI) reconstruction and synthesis are becoming central to embodied AI and AR/VR. Yet, despite rapid progress, existing HOI generation research remains fragmented across three disjoint tracks: (1) pose-only synthesis…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Mingju Gao , Kaisen Yang , Huan-ang Gao , Bohan Li , Ao Ding , Wenyi Li , Yangcheng Yu , Jinkun Liu , Shaocong Xu , Yike Niu , Haohan Chi , Hao Chen , Hao Tang , Yu Zhang , Li Yi , Hao Zhao

Recent vision-language-action (VLA) models can generate plausible end-effector motions, yet they often fail in long-horizon, contact-rich tasks because the underlying hand-object interaction (HOI) structure is not explicitly represented. An…

Robotics · Computer Science 2026-02-17 Huajian Zeng , Lingyun Chen , Jiaqi Yang , Yuantai Zhang , Fan Shi , Peidong Liu , Xingxing Zuo

Human-centric video generation has advanced rapidly, yet existing methods struggle to produce controllable and physically consistent Human-Object Interaction (HOI) videos. Existing works rely on dense control signals, template videos, or…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Jiazhi Guan , Quanwei Yang , Luying Huang , Junhao Liang , Borong Liang , Haocheng Feng , Wei He , Kaisiyuan Wang , Hang Zhou , Jingdong Wang

Joint reconstruction of 3D human and object from a single image is an active research area, with pivotal applications in robotics and digital content creation. Despite recent advances, existing approaches suffer from two fundamental…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Hyeongjin Nam , Daniel Sungho Jung , Kyoung Mu Lee

Generating natural hand-object interactions in 3D is challenging as the resulting hand and object motions are expected to be physically plausible and semantically meaningful. Furthermore, generalization to unseen objects is hindered by the…

Computer Vision and Pattern Recognition · Computer Science 2024-12-24 Sammy Christen , Shreyas Hampali , Fadime Sener , Edoardo Remelli , Tomas Hodan , Eric Sauser , Shugao Ma , Bugra Tekin

Generating talking avatars is a fundamental task in video generation. Although existing methods can generate full-body talking avatars with simple human motion, extending this task to grounded human-object interaction (GHOI) remains an open…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Youliang Zhang , Zhengguang Zhou , Zhentao Yu , Ziyao Huang , Teng Hu , Sen Liang , Guozhen Zhang , Ziqiao Peng , Shunkai Li , Yi Chen , Zixiang Zhou , Yuan Zhou , Qinglin Lu , Xiu Li

Learning human-object manipulation presents significant challenges due to its fine-grained and contact-rich nature of the motions involved. Traditional physics-based animation requires extensive modeling and manual setup, and more…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Quankai Gao , Jiawei Yang , Qiangeng Xu , Le Chen , Yue Wang

This paper addresses new methodologies to deal with the challenging task of generating dynamic Human-Object Interactions from textual descriptions (Text2HOI). While most existing works assume interactions with limited body parts or static…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Qianyang Wu , Ye Shi , Xiaoshui Huang , Jingyi Yu , Lan Xu , Jingya Wang

Functional grasp is essential for enabling dexterous multi-finger robot hands to manipulate objects effectively. However, most prior work either focuses on power grasping, which simply involves holding an object still, or relies on costly…

Computer Vision and Pattern Recognition · Computer Science 2025-05-14 Hongyi Chen , Yunchao Yao , Yufei Ye , Zhixuan Xu , Homanga Bharadhwaj , Jiashun Wang , Shubham Tulsiani , Zackory Erickson , Jeffrey Ichnowski

Generating realistic 3D Human-Object Interactions (HOI) is a fundamental task for applications ranging from embodied AI to virtual content creation, which requires harmonizing high-level semantic intent with strict low-level physical…

Robotics · Computer Science 2026-05-08 Hao Wang , Shiqi Wang , Qi Liu

We present a system that transforms speech into physical objects using 3D generative AI and discrete robotic assembly. By leveraging natural language, the system makes design and manufacturing more accessible to people without expertise in…

Robotics · Computer Science 2025-09-17 Alexander Htet Kyaw , Miana Smith , Se Hwan Jeon , Neil Gershenfeld

Synthesizing 3D human avatars interacting realistically with a scene is an important problem with applications in AR/VR, video games and robotics. Towards this goal, we address the task of generating a virtual human -- hands and full body…

Robotics · Computer Science 2023-03-30 Purva Tendulkar , Dídac Surís , Carl Vondrick

We address the problem of generating diverse 3D human motions from textual descriptions. This challenging task requires joint modeling of both modalities: understanding and extracting useful human-centric information from the text, and then…

Computer Vision and Pattern Recognition · Computer Science 2022-07-25 Mathis Petrovich , Michael J. Black , Gül Varol

Generating human-object interactions (HOIs) is critical with the tremendous advances of digital avatars. Existing datasets are typically limited to humans interacting with a single object while neglecting the ubiquitous manipulation of…

Computer Vision and Pattern Recognition · Computer Science 2024-09-12 Xintao Lv , Liang Xu , Yichao Yan , Xin Jin , Congsheng Xu , Shuwen Wu , Yifan Liu , Lincheng Li , Mengxiao Bi , Wenjun Zeng , Xiaokang Yang

Compositional Customized Image Generation aims to customize multiple target concepts within generation content, which has gained attention for its wild application. Existing approaches mainly concentrate on the target entity's appearance…

Computer Vision and Pattern Recognition · Computer Science 2025-08-29 Zhu Xu , Zhaowen Wang , Yuxin Peng , Yang Liu

Recent advancements in generative models have enabled the creation of dynamic 4D content - 3D objects in motion - based on text prompts, which holds potential for applications in virtual worlds, media, and gaming. Existing methods provide…

Computer Vision and Pattern Recognition · Computer Science 2025-05-28 Ohad Rahamim , Ori Malca , Dvir Samuel , Gal Chechik

Synthesizing physically plausible articulated human-object interactions (HOI) without 3D/4D supervision remains a fundamental challenge. While recent zero-shot approaches leverage video diffusion models to synthesize human-object…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Zihao Huang , Tianqi Liu , Zhaoxi Chen , Shaocong Xu , Saining Zhang , Lixing Xiao , Zhiguo Cao , Wei Li , Hao Zhao , Ziwei Liu

Text-to-video (T2V) generation has made tremendous progress in generating complicated scenes based on texts. However, human-object interaction (HOI) often cannot be precisely generated by current T2V models due to the lack of large-scale…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Kun Liu , Qi Liu , Xinchen Liu , Jie Li , Yongdong Zhang , Jiebo Luo , Xiaodong He , Wu Liu

Text-to-3D generation is a valuable technology in virtual reality and digital content creation. While recent works have pushed the boundaries of text-to-3D generation, producing high-fidelity 3D objects with inefficient prompts and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Wenqing Wang , Yun Fu