English
Related papers

Related papers: Follow My Hold: Hand-Object Interaction Reconstruc…

200 papers

Predicting and generating human hand grasp over objects is critical for animation and robotic tasks. In this work, we focus on generating both the hand and objects in a grasp by a single diffusion model. Our proposed Joint Hand-Object…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 Jinkun Cao , Jingyuan Liu , Kris Kitani , Yi Zhou

Object-level manipulation, relocating or reorienting objects in images or videos while preserving scene realism, is central to film post-production, AR, and creative editing. Yet existing methods struggle to jointly achieve three core…

Computer Vision and Pattern Recognition · Computer Science 2026-02-13 Penghui Ruan , Bojia Zi , Xianbiao Qi , Youze Huang , Rong Xiao , Pichao Wang , Jiannong Cao , Yuhui Shi

This paper presents an approach that reconstructs a hand-held object from a monocular video. In contrast to many recent methods that directly predict object geometry by a trained network, the proposed approach does not require any learned…

Computer Vision and Pattern Recognition · Computer Science 2022-12-01 Di Huang , Xiaopeng Ji , Xingyi He , Jiaming Sun , Tong He , Qing Shuai , Wanli Ouyang , Xiaowei Zhou

Understanding how humans would behave during hand-object interaction is vital for applications in service robot manipulation and extended reality. To achieve this, some recent works have been proposed to simultaneously forecast hand…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Junyi Ma , Jingyi Xu , Xieyuanli Chen , Hesheng Wang

Our work aims to obtain 3D reconstruction of hands and manipulated objects from monocular videos. Reconstructing hand-object manipulations holds a great potential for robotics and learning from human demonstrations. The supervised learning…

Computer Vision and Pattern Recognition · Computer Science 2022-03-15 Yana Hasson , Gül Varol , Ivan Laptev , Cordelia Schmid

Our work aims to reconstruct a 3D object that is held and rotated by a hand in front of a static RGB camera. Previous methods that use implicit neural representations to recover the geometry of a generic hand-held object from multi-view…

Computer Vision and Pattern Recognition · Computer Science 2023-12-29 Shijian Jiang , Qi Ye , Rengan Xie , Yuchi Huo , Xiang Li , Yang Zhou , Jiming Chen

This paper focuses on a challenging setting of simultaneously modeling geometry and appearance of hand-object interaction scenes without any object priors. We follow the trend of dynamic 3D Gaussian Splatting based methods, and address…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Hao Tian , Chenyangguang Zhang , Rui Liu , Wen Shen , Xiaolin Qin

Most RGB-based hand-object reconstruction methods rely on object templates, while template-free methods typically assume full object visibility. This assumption often breaks in real-world settings, where fixed camera viewpoints and static…

Computer Vision and Pattern Recognition · Computer Science 2025-08-08 Shibo Wang , Haonan He , Maria Parelli , Christoph Gebhardt , Zicong Fan , Jie Song

We present GSD, a diffusion model approach based on Gaussian Splatting (GS) representation for 3D object reconstruction from a single view. Prior works suffer from inconsistent 3D geometry or mediocre rendering quality due to improper…

Computer Vision and Pattern Recognition · Computer Science 2024-10-30 Yuxuan Mu , Xinxin Zuo , Chuan Guo , Yilin Wang , Juwei Lu , Xiaofeng Wu , Songcen Xu , Peng Dai , Youliang Yan , Li Cheng

This work presents a flexible system to reconstruct 3D models of objects captured with an RGB-D sensor. A major advantage of the method is that our reconstruction pipeline allows the user to acquire a full 3D model of the object. This is…

Computer Vision and Pattern Recognition · Computer Science 2015-05-22 Aitor Aldoma , Johann Prankl , Alexander Svejda , Markus Vincze

Occlusion is a common issue in 3D reconstruction from RGB-D videos, often blocking the complete reconstruction of objects and presenting an ongoing problem. In this paper, we propose a novel framework, empowered by a 2D diffusion-based…

Computer Vision and Pattern Recognition · Computer Science 2024-03-20 Yubin Hu , Sheng Ye , Wang Zhao , Matthieu Lin , Yuze He , Yu-Hui Wen , Ying He , Yong-Jin Liu

Recent generative models can synthesize high-quality images, but they often fail to generate humans interacting with objects using their hands. This arises mostly from the model's misunderstanding of such interactions and the hardships of…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Patrick Kwon , Chen Chen , Hanbyul Joo

Existing reconstruction or hand-object pose estimation methods are capable of producing coarse interaction states. However, due to the complex and diverse geometry of both human hands and objects, these approaches often suffer from…

Computer Vision and Pattern Recognition · Computer Science 2025-09-04 Miao Xu , Xiangyu Zhu , Xusheng Liang , Zidu Wang , Jinlin Wu , Zhen Lei

Our work aims to reconstruct hand-held objects given a single RGB image. In contrast to prior works that typically assume known 3D templates and reduce the problem to 3D pose estimation, our work reconstructs generic hand-held object…

Computer Vision and Pattern Recognition · Computer Science 2022-04-15 Yufei Ye , Abhinav Gupta , Shubham Tulsiani

Joint reconstruction of human-object interaction marks a significant milestone in comprehending the intricate interrelations between humans and their surrounding environment. Nevertheless, previous optimization methods often struggle to…

Computer Vision and Pattern Recognition · Computer Science 2025-09-10 Ao Li , Jinpeng Liu , Yixuan Zhu , Yansong Tang

Diffusion models have made breakthroughs in 3D generation tasks. Current 3D diffusion models focus on reconstructing target shape from images or a set of partial observations. While excelling in global context understanding, they struggle…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Yuanbo Wang , Zhaoxuan Zhang , Jiajin Qiu , Dilong Sun , Zhengyu Meng , Xiaopeng Wei , Xin Yang

The correct insertion of virtual objects in images of real-world scenes requires a deep understanding of the scene's lighting, geometry and materials, as well as the image formation process. While recent large-scale diffusion models have…

Computer Vision and Pattern Recognition · Computer Science 2024-08-20 Ruofan Liang , Zan Gojcic , Merlin Nimier-David , David Acuna , Nandita Vijaykumar , Sanja Fidler , Zian Wang

Reconstructing compositional 3D representations of scenes, where each object is represented with its own 3D model, is a highly desirable capability in robotics and augmented reality. However, most existing methods rely heavily on strong…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Vincent van der Brugge , Marc Pollefeys , Joshua B. Tenenbaum , Ayush Tewari , Krishna Murthy Jatavallabhula

Hand pose estimation from a single image has many applications. However, approaches to full 3D body pose estimation are typically trained on day-to-day activities or actions. As such, detailed hand-to-hand interactions are poorly…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Maksym Ivashechkin , Oscar Mendez , Richard Bowden

Reconstructing a renderable 3D model from images is a useful but challenging task. Recent feedforward 3D reconstruction methods have demonstrated remarkable success in efficiently recovering geometry, but still cannot accurately model the…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Zekai Gu , Shuoxuan Feng , Yansong Wang , Hanzhuo Huang , Zhongshuo Du , Chengfeng Zhao , Chengwei Ren , Peng Wang , Yuan Liu