English
Related papers

Related papers: ForeHOI: Feed-forward 3D Object Reconstruction fro…

200 papers

Existing methods for 3D tracking from monocular RGB videos predominantly consider articulated and rigid objects. Modelling dense non-rigid object deformations in this setting remained largely unaddressed so far, although such effects can…

Computer Vision and Pattern Recognition · Computer Science 2023-10-16 Soshi Shimada , Vladislav Golyanik , Patrick Pérez , Christian Theobalt

Existing methods for reconstructing objects and humans from a monocular image suffer from severe mesh collisions and performance limitations for interacting occluding objects. This paper introduces a method to obtain a globally consistent…

Computer Vision and Pattern Recognition · Computer Science 2024-08-16 Sarthak Batra , Partha P. Chakrabarti , Simon Hadfield , Armin Mustafa

Monocular 3D hand reconstruction is intrinsically a geometric problem, yet RGB appearance features alone often struggle to resolve severe ambiguities caused by self-occlusions and hand-object interactions. While introducing depth can…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Weiquan Lin , Yaoqing Hu , Liangchen Dai , Xu Tang , Xingyu Chen

We introduce UniCon3R, a unified feed-forward framework for online human-scene 4D reconstruction from monocular video. Current feed-forward human-scene reconstruction methods suffer from artifacts, where bodies float above the ground or…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Tanuj Sur , Shashank Tripathi , Nikos Athanasiou , Ha Linh Nguyen , Kai Xu , Michael J. Black , Angela Yao

Recovering 3D Human-Object Interaction (HOI) from single color images is challenging due to depth ambiguities, occlusions, and the huge variation in object shape and appearance. Thus, past work requires controlled settings such as known…

Computer Vision and Pattern Recognition · Computer Science 2025-04-25 Alpár Cseke , Shashank Tripathi , Sai Kumar Dwivedi , Arjun Lakshmipathy , Agniv Chatterjee , Michael J. Black , Dimitrios Tzionas

We address the challenging task of anticipating human-object interaction in first person videos. Most existing methods ignore how the camera wearer interacts with the objects, or simply consider body motion as a separate modality. In…

Computer Vision and Pattern Recognition · Computer Science 2020-07-21 Miao Liu , Siyu Tang , Yin Li , James Rehg

We introduce a novel framework for reconstructing dynamic human-object interactions from monocular video that overcomes challenges associated with occlusions and temporal inconsistencies. Traditional 3D reconstruction methods typically…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Hyungjun Doh , Dong In Lee , Seunggeun Chi , Pin-Hao Huang , Kwonjoon Lee , Sangpil Kim , Karthik Ramani

We present a new framework to reconstruct holistic 3D indoor scenes including both room background and indoor objects from single-view images. Existing methods can only produce 3D shapes of indoor objects with limited geometry quality…

Computer Vision and Pattern Recognition · Computer Science 2022-08-09 Haolin Liu , Yujian Zheng , Guanying Chen , Shuguang Cui , Xiaoguang Han

Synthesizing realistic human-object interactions (HOI) in video is challenging due to the complex, instance-specific interaction dynamics of both humans and objects. Incorporating controllability in video generation further adds to the…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Wanyue Zhang , Lin Geng Foo , Thabo Beeler , Rishabh Dabral , Christian Theobalt

High-fidelity 3D scene reconstruction from monocular videos continues to be challenging, especially for complete and fine-grained geometry reconstruction. The previous 3D reconstruction approaches with neural implicit representations have…

Computer Vision and Pattern Recognition · Computer Science 2022-10-03 Zi-Xin Zou , Shi-Sheng Huang , Yan-Pei Cao , Tai-Jiang Mu , Ying Shan , Hongbo Fu

Human-object interaction detection (HOID) refers to localizing interactive human-object pairs in images and identifying the interactions. Since there could be an exponential number of object-action combinations, labeled data is limited -…

Computer Vision and Pattern Recognition · Computer Science 2025-07-18 Sandipan Sarma , Agney Talwarr , Arijit Sur

We propose CG-HOI, the first method to address the task of generating dynamic 3D human-object interactions (HOIs) from text. We model the motion of both human and object in an interdependent fashion, as semantically rich human motion rarely…

Computer Vision and Pattern Recognition · Computer Science 2024-05-20 Christian Diller , Angela Dai

Prior works for reconstructing hand-held objects from a single image train models on images paired with 3D shapes. Such data is challenging to gather in the real world at scale. Consequently, these approaches do not generalize well when…

Computer Vision and Pattern Recognition · Computer Science 2024-09-24 Aditya Prakash , Matthew Chang , Matthew Jin , Ruisen Tu , Saurabh Gupta

Monocular dynamic reconstruction is a challenging and long-standing vision problem due to the highly ill-posed nature of the task. Existing approaches depend on templates, are effective only in quasi-static scenes, or fail to model 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Qianqian Wang , Vickie Ye , Hang Gao , Weijia Zeng , Jake Austin , Zhengqi Li , Angjoo Kanazawa

We present a novel method for monocular hand shape and pose estimation at unprecedented runtime performance of 100fps and at state-of-the-art accuracy. This is enabled by a new learning based architecture designed such that it can make use…

Computer Vision and Pattern Recognition · Computer Science 2022-03-14 Yuxiao Zhou , Marc Habermann , Weipeng Xu , Ikhsanul Habibie , Christian Theobalt , Feng Xu

We propose Mesh4D, a feed-forward model for monocular 4D mesh reconstruction. Given a monocular video of a dynamic object, our model reconstructs the object's complete 3D shape and motion, represented as a deformation field. Our key…

Computer Vision and Pattern Recognition · Computer Science 2026-01-09 Zeren Jiang , Chuanxia Zheng , Iro Laina , Diane Larlus , Andrea Vedaldi

Monocular image-based 3D reconstruction of faces is a long-standing problem in computer vision. Since image data is a 2D projection of a 3D face, the resulting depth ambiguity makes the problem ill-posed. Most existing methods rely on…

Computer Vision and Pattern Recognition · Computer Science 2019-04-10 Ayush Tewari , Florian Bernard , Pablo Garrido , Gaurav Bharaj , Mohamed Elgharib , Hans-Peter Seidel , Patrick Pérez , Michael Zollhöfer , Christian Theobalt

Reconstructing the absolute 3D pose and shape of the hands from the user's viewpoint using a single head-mounted camera is crucial for practical egocentric interaction in AR/VR, telepresence, and hand-centric manipulation tasks, where…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Christen Millerdurai , Shaoxiang Wang , Yaxu Xie , Vladislav Golyanik , Didier Stricker , Alain Pagani

Fine-grained capturing of 3D HOI boosts human activity understanding and facilitates downstream visual tasks, including action recognition, holistic scene reconstruction, and human motion synthesis. Despite its significance, existing works…

Computer Vision and Pattern Recognition · Computer Science 2023-12-19 Nan Jiang , Tengyu Liu , Zhexuan Cao , Jieming Cui , Zhiyuan zhang , Yixin Chen , He Wang , Yixin Zhu , Siyuan Huang

We present the first approach to volumetric performance capture and novel-view rendering at real-time speed from monocular video, eliminating the need for expensive multi-view systems or cumbersome pre-acquisition of a personalized template…

Computer Vision and Pattern Recognition · Computer Science 2020-07-29 Ruilong Li , Yuliang Xiu , Shunsuke Saito , Zeng Huang , Kyle Olszewski , Hao Li