English
Related papers

Related papers: HOIDiffusion: Generating Realistic 3D Hand-Object …

200 papers

Humans frequently grasp, manipulate, and move objects. Interactive systems assist humans in these tasks, enabling applications in Embodied AI, human-robot interaction, and virtual reality. However, current methods in hand-object synthesis…

Robotics · Computer Science 2025-03-10 Sammy Christen

We introduce D3D-HOI: a dataset of monocular videos with ground truth annotations of 3D object pose, shape and part motion during human-object interactions. Our dataset consists of several common articulated objects captured from diverse…

Computer Vision and Pattern Recognition · Computer Science 2021-08-20 Xiang Xu , Hanbyul Joo , Greg Mori , Manolis Savva

Generating 3D scenes from human motion sequences supports numerous applications, including virtual reality and architectural design. However, previous auto-regression-based human-aware 3D scene generation methods have struggled to…

Computer Vision and Pattern Recognition · Computer Science 2024-08-21 Xiaolin Hong , Hongwei Yi , Fazhi He , Qiong Cao

Supervised learning models for precise tracking of hand-object interactions (HOI) in 3D require large amounts of annotated data for training. Moreover, it is not intuitive for non-experts to label 3D ground truth (e.g. 6DoF object pose) on…

Computer Vision and Pattern Recognition · Computer Science 2024-02-05 Chengyan Zhang , Rahul Chaudhari

This paper explores a cross-modality synthesis task that infers 3D human-object interactions (HOIs) from a given text-based instruction. Existing text-to-HOI synthesis methods mainly deploy a direct mapping from texts to object-specific 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-03-05 Xuehao Gao , Yang Yang , Shaoyi Du , Yang Wu , Yebin Liu , Guo-Jun Qi

Generating realistic 3D scenes is an area of growing interest in computer vision and robotics. However, creating high-quality, diverse synthetic 3D content often requires expert intervention, making it costly and complex. Recently, efforts…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Siyi Hu , Diego Martin Arroyo , Stephanie Debats , Fabian Manhardt , Luca Carlone , Federico Tombari

Reconstructing human-object interaction in 3D from a single RGB image is a challenging task and existing data driven methods do not generalize beyond the objects present in the carefully curated 3D interaction datasets. Capturing…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Xianghui Xie , Bharat Lal Bhatnagar , Jan Eric Lenssen , Gerard Pons-Moll

To serve as a scalable data source for embodied AI, world models should act as true simulators that infer interaction dynamics strictly from user actions, rather than mere conditional video generators relying on privileged future object…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Dayou Li , Lulin Liu , Bangya Liu , Shijie Zhou , Jiu Feng , Ziqi Lu , Minghui Zheng , Chenyu You , Zhiwen Fan

Existing reconstruction or hand-object pose estimation methods are capable of producing coarse interaction states. However, due to the complex and diverse geometry of both human hands and objects, these approaches often suffer from…

Computer Vision and Pattern Recognition · Computer Science 2025-09-04 Miao Xu , Xiangyu Zhu , Xusheng Liang , Zidu Wang , Jinlin Wu , Zhen Lei

Human-Object Interaction (HOI) detection lies at the core of action understanding. Besides 2D information such as human/object appearance and locations, 3D pose is also usually utilized in HOI learning since its view-independence. However,…

Computer Vision and Pattern Recognition · Computer Science 2020-05-22 Yong-Lu Li , Xinpeng Liu , Han Lu , Shiyi Wang , Junqi Liu , Jiefeng Li , Cewu Lu

How can we reconstruct 3D hand poses when large portions of the hand are heavily occluded by itself or by objects? Humans often resolve such ambiguities by leveraging contextual knowledge -- such as affordances, where an object's shape and…

Computer Vision and Pattern Recognition · Computer Science 2025-10-02 Naru Suzuki , Takehiko Ohkawa , Tatsuro Banno , Jihyun Lee , Ryosuke Furuta , Yoichi Sato

We present HuGDiffusion, a generalizable 3D Gaussian splatting (3DGS) learning pipeline to achieve novel view synthesis (NVS) of human characters from single-view input images. Existing approaches typically require monocular videos or…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Yingzhi Tang , Qijian Zhang , Junhui Hou

Understanding how humans interact with the surrounding environment, and specifically reasoning about object interactions and affordances, is a critical challenge in computer vision, robotics, and AI. Current approaches often depend on…

Computer Vision and Pattern Recognition · Computer Science 2026-02-12 Harry Zhang , Luca Carlone

3D human generation is an important problem with a wide range of applications in computer vision and graphics. Despite recent progress in generative AI such as diffusion models or rendering methods like Neural Radiance Fields or Gaussian…

Computer Vision and Pattern Recognition · Computer Science 2025-06-06 Maksym Ivashechkin , Oscar Mendez , Richard Bowden

Recent text-to-image generative models have exhibited remarkable abilities in generating high-fidelity and photo-realistic images. However, despite the visually impressive results, these models often struggle to preserve plausible human…

Computer Vision and Pattern Recognition · Computer Science 2024-01-02 Zhenzhen Weng , Laura Bravo-Sánchez , Serena Yeung-Levy

While diffusion models and large-scale motion datasets have advanced text-driven human motion synthesis, extending these advances to 4D human-object interaction (HOI) remains challenging, mainly due to the limited availability of…

Computer Vision and Pattern Recognition · Computer Science 2025-06-19 Shujia Li , Haiyu Zhang , Xinyuan Chen , Yaohui Wang , Yutong Ban

Modeling hand-object manipulations is essential for understanding how humans interact with their environment. While of practical importance, estimating the pose of hands and objects during interactions is challenging due to the large mutual…

Computer Vision and Pattern Recognition · Computer Science 2020-04-29 Yana Hasson , Bugra Tekin , Federica Bogo , Ivan Laptev , Marc Pollefeys , Cordelia Schmid

We tackle the task of reconstructing hand-object interactions from short video clips. Given an input video, our approach casts 3D inference as a per-video optimization and recovers a neural 3D representation of the object shape, as well as…

Computer Vision and Pattern Recognition · Computer Science 2023-09-12 Yufei Ye , Poorvi Hebbar , Abhinav Gupta , Shubham Tulsiani

Understanding hand-object interaction (HOI) is fundamental to computer vision, robotics, and AR/VR. However, conventional hand videos often lack essential physical information such as contact forces and motion signals, and are prone to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Xinyu Zhang , Ziyi Kou , Chuan Qin , Mia Huang , Ergys Ristani , Ankit Kumar , Lele Chen , Kun He , Abdeslam Boularias , Li Guan

While large-scale human motion capture datasets have advanced human motion generation, modeling and generating dynamic 3D human-object interactions (HOIs) remain challenging due to dataset limitations. Existing datasets often lack…

Computer Vision and Pattern Recognition · Computer Science 2025-09-12 Sirui Xu , Dongting Li , Yucheng Zhang , Xiyan Xu , Qi Long , Ziyin Wang , Yunzhi Lu , Shuchang Dong , Hezi Jiang , Akshat Gupta , Yu-Xiong Wang , Liang-Yan Gui