English
Related papers

Related papers: Reconstructing Hand-Held Objects in 3D from Images…

200 papers

Recent approaches to jointly reconstruct 3D humans and objects from a single RGB image represent 3D shapes with template-based or coarse models, which fail to capture details of loose clothing on human bodies. In this paper, we introduce a…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Ayushi Dutta , Marco Pesavento , Marco Volino , Adrian Hilton , Armin Mustafa

Existing hand-object interactions (HOI) methods are largely limited to rigid objects, while 4D reconstruction methods of articulated objects generally require pre-scanning the object or even multi-view videos. It remains an unexplored but…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Zikai Wang , Zhilu Zhang , Yiqing Wang , Hui Li , Wangmeng Zuo

Monocular 3D hand mesh recovery is challenging due to high degrees of freedom of hands, 2D-to-3D ambiguity and self-occlusion. Most existing methods are either inefficient or less straightforward for predicting the position of 3D mesh…

Computer Vision and Pattern Recognition · Computer Science 2025-12-08 Yihong Lin , Xianjia Wu , Xilai Wang , Jianqiao Hu , Songju Lei , Xiandong Li , Wenxiong Kang

Semantic aware reconstruction is more advantageous than geometric-only reconstruction for future robotic and AR/VR applications because it represents not only where things are, but also what things are. Object-centric mapping is a task to…

Computer Vision and Pattern Recognition · Computer Science 2021-02-16 Kejie Li , Hamid Rezatofighi , Ian Reid

Existing methods for 3D tracking from monocular RGB videos predominantly consider articulated and rigid objects. Modelling dense non-rigid object deformations in this setting remained largely unaddressed so far, although such effects can…

Computer Vision and Pattern Recognition · Computer Science 2023-10-16 Soshi Shimada , Vladislav Golyanik , Patrick Pérez , Christian Theobalt

Monocular 3D hand reconstruction is intrinsically a geometric problem, yet RGB appearance features alone often struggle to resolve severe ambiguities caused by self-occlusions and hand-object interactions. While introducing depth can…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Weiquan Lin , Yaoqing Hu , Liangchen Dai , Xu Tang , Xingyu Chen

3D hand-mesh reconstruction from RGB images facilitates many applications, including augmented reality (AR). However, this requires not only real-time speed and accurate hand pose and shape but also plausible mesh-image alignment. While…

Computer Vision and Pattern Recognition · Computer Science 2021-09-07 Xiao Tang , Tianyu Wang , Chi-Wing Fu

We develop a system for modeling hand-object interactions in 3D from RGB images that show a hand which is holding a novel object from a known category. We design a Convolutional Neural Network (CNN) for Hand-held Object Pose and Shape…

Computer Vision and Pattern Recognition · Computer Science 2019-11-12 Mia Kokic , Danica Kragic , Jeannette Bohg

Reconstructing the motion of objects from videos is a key component for embodied AI and robot manipulation. While diverse approaches to object pose tracking have been studied, they rely heavily on strong external priors, such as depth data…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Jisu Shin , Junoh Lee , JunGyu Lee , Inhwan Bae , Dohyeon Lee , Hokyun Im , Youngwoon Lee , Hae-Gon Jeon

We introduce a data capture system and a new dataset, HO-Cap, for 3D reconstruction and pose tracking of hands and objects in videos. The system leverages multiple RGBD cameras and a HoloLens headset for data collection, avoiding the use of…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Jikai Wang , Qifan Zhang , Yu-Wei Chao , Bowen Wen , Xiaohu Guo , Yu Xiang

Reconstructing the hand mesh from one single RGB image is a challenging task because hands are often occluded by other objects. Most previous works attempt to explore more additional information and adopt attention mechanisms for improving…

Computer Vision and Pattern Recognition · Computer Science 2025-06-16 Zixun Jiao , Xihan Wang , Zhaoqiang Xia , Lianhe Shao , Quanli Gao

We ask whether everyday open-world monocular videos can be turned into reusable 4D interaction primitives: articulated hand motion, object shape with 6D pose over time, and the when/where of contact. Such a capability would enable scalable…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Hao Xu , Yilin Liu , Yinqiao Wang , Chi-Wing Fu , Niloy J. Mitra

We introduce the task of Reconstructing Objects along Hand Interaction Timelines (ROHIT). We first define the Hand Interaction Timeline (HIT) from a rigid object's perspective. In a HIT, an object is first static relative to the scene, then…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Zhifan Zhu , Siddhant Bansal , Shashank Tripathi , Dima Damen

Joint reconstruction of 3D human and object from a single image is an active research area, with pivotal applications in robotics and digital content creation. Despite recent advances, existing approaches suffer from two fundamental…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Hyeongjin Nam , Daniel Sungho Jung , Kyoung Mu Lee

Learning-based 3D object reconstruction enables single- or few-shot estimation of 3D object models. For robotics, this holds the potential to allow model-based methods to rapidly adapt to novel objects and scenes. Existing 3D reconstruction…

Recent advancements in 3D object reconstruction from single images have primarily focused on improving the accuracy of object shapes. Yet, these techniques often fail to accurately capture the inter-relation between the object, ground, and…

Computer Vision and Pattern Recognition · Computer Science 2024-07-29 Yunze Man , Yichen Sheng , Jianming Zhang , Liang-Yan Gui , Yu-Xiong Wang

Physical contact provides additional constraints for hand-object state reconstruction as well as a basis for further understanding of interaction affordances. Estimating these severely occluded regions from monocular images presents a…

Computer Vision and Pattern Recognition · Computer Science 2022-05-03 Zimeng Zhao , Binghui Zuo , Wei Xie , Yangang Wang

We interact with the world with our hands and see it through our own (egocentric) perspective. A holistic 3Dunderstanding of such interactions from egocentric views is important for tasks in robotics, AR/VR, action recognition and motion…

Monocular 3D reconstruction of deformable objects, such as human body parts, has been typically approached by predicting parameters of heavyweight linear models. In this paper, we demonstrate an alternative solution that is based on the…

Computer Vision and Pattern Recognition · Computer Science 2019-08-06 Dominik Kulon , Haoyang Wang , Riza Alp Güler , Michael Bronstein , Stefanos Zafeiriou

Narrated instructional videos often show and describe manipulations of similar objects, e.g., repairing a particular model of a car or laptop. In this work we aim to reconstruct such objects and to localize associated narrations in 3D.…

Computer Vision and Pattern Recognition · Computer Science 2021-09-13 Dimitri Zhukov , Ignacio Rocco , Ivan Laptev , Josef Sivic , Johannes L. Schönberger , Bugra Tekin , Marc Pollefeys