English
Related papers

Related papers: RHINO: Reconstructing Human Interactions with Nove…

200 papers

We ask whether everyday open-world monocular videos can be turned into reusable 4D interaction primitives: articulated hand motion, object shape with 6D pose over time, and the when/where of contact. Such a capability would enable scalable…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Hao Xu , Yilin Liu , Yinqiao Wang , Chi-Wing Fu , Niloy J. Mitra

The reconstruction of three-dimensional dynamic scenes is a well-established yet challenging task within the domain of computer vision. In this paper, we propose a novel approach that combines the domains of 3D geometry reconstruction and…

Computer Vision and Pattern Recognition · Computer Science 2025-09-11 David Stotko , Reinhard Klein

Since humans interact with diverse objects every day, the holistic 3D capture of these interactions is important to understand and model human behaviour. However, most existing methods for hand-object reconstruction from RGB either assume…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Zicong Fan , Maria Parelli , Maria Eleni Kadoglou , Muhammed Kocabas , Xu Chen , Michael J. Black , Otmar Hilliges

This paper presents an approach that reconstructs a hand-held object from a monocular video. In contrast to many recent methods that directly predict object geometry by a trained network, the proposed approach does not require any learned…

Computer Vision and Pattern Recognition · Computer Science 2022-12-01 Di Huang , Xiaopeng Ji , Xingyi He , Jiaming Sun , Tong He , Qing Shuai , Wanli Ouyang , Xiaowei Zhou

Existing methods for 3D tracking from monocular RGB videos predominantly consider articulated and rigid objects. Modelling dense non-rigid object deformations in this setting remained largely unaddressed so far, although such effects can…

Computer Vision and Pattern Recognition · Computer Science 2023-10-16 Soshi Shimada , Vladislav Golyanik , Patrick Pérez , Christian Theobalt

Reconstructing 3D models of dynamic, real-world objects with high-fidelity textures from monocular frame sequences has been a challenging problem in recent years. This difficulty stems from factors such as shadows, indirect illumination,…

Computer Vision and Pattern Recognition · Computer Science 2025-05-12 Alakh Aggarwal , Ningna Wang , Xiaohu Guo

Reconstructing hand-held objects from monocular RGB images is an appealing yet challenging task. In this task, contacts between hands and objects provide important cues for recovering the 3D geometry of the hand-held objects. Though recent…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Junxing Hu , Hongwen Zhang , Zerui Chen , Mengcheng Li , Yunlong Wang , Yebin Liu , Zhenan Sun

We present a method to reconstruct time-consistent human body models from monocular videos, focusing on extremely loose clothing or handheld object interactions. Prior work in human reconstruction is either limited to tight clothing with no…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Jeff Tan , Donglai Xiang , Shubham Tulsiani , Deva Ramanan , Gengshan Yang

Demystifying complex human-ground interactions is essential for accurate and realistic 3D human motion reconstruction from RGB videos, as it ensures consistency between the humans and the ground plane. Prior methods have modeled…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Sihan Ma , Qiong Cao , Hongwei Yi , Jing Zhang , Dacheng Tao

We propose RoHM, an approach for robust 3D human motion reconstruction from monocular RGB(-D) videos in the presence of noise and occlusions. Most previous approaches either train neural networks to directly regress motion in 3D or learn…

Computer Vision and Pattern Recognition · Computer Science 2024-04-16 Siwei Zhang , Bharat Lal Bhatnagar , Yuanlu Xu , Alexander Winkler , Petr Kadlecek , Siyu Tang , Federica Bogo

We propose a method for generating video-realistic animations of real humans under user control. In contrast to conventional human character rendering, we do not require the availability of a production-quality photo-realistic 3D model of…

Computer Vision and Pattern Recognition · Computer Science 2019-05-13 Lingjie Liu , Weipeng Xu , Michael Zollhoefer , Hyeongwoo Kim , Florian Bernard , Marc Habermann , Wenping Wang , Christian Theobalt

This paper proposes GraviCap, i.e., a new approach for joint markerless 3D human motion capture and object trajectory estimation from monocular RGB videos. We focus on scenes with objects partially observed during a free flight. In contrast…

Computer Vision and Pattern Recognition · Computer Science 2021-08-20 Rishabh Dabral , Soshi Shimada , Arjun Jain , Christian Theobalt , Vladislav Golyanik

Character video synthesis aims to produce realistic videos of animatable characters within lifelike scenes. As a fundamental problem in the computer vision and graphics community, 3D works typically require multi-view captures for per-case…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Yifang Men , Yuan Yao , Miaomiao Cui , Liefeng Bo

In this paper, we address the challenge of reconstructing general articulated 3D objects from a single video. Existing works employing dynamic neural radiance fields have advanced the modeling of articulated objects like humans and animals…

Computer Vision and Pattern Recognition · Computer Science 2024-04-18 Chaoyue Song , Jiacheng Wei , Chuan-Sheng Foo , Guosheng Lin , Fayao Liu

We introduce a novel, data-driven approach for reconstructing temporally coherent 3D motion from unstructured and potentially partial observations of non-rigidly deforming shapes. Our goal is to achieve high-fidelity motion reconstructions…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Aymen Merrouche , Stefanie Wuhrer , Edmond Boyer

The intricate nature of real-world driving environments, characterized by dynamic and diverse interactions among multiple vehicles and their possible future states, presents considerable challenges in accurately predicting the motion states…

Robotics · Computer Science 2025-08-13 Keshu Wu , Yang Zhou , Haotian Shi , Dominique Lord , Bin Ran , Xinyue Ye

Monocular 3D clothed human reconstruction aims to create a complete 3D avatar from a single image. To tackle the human geometry lacking in one RGB image, current methods typically resort to a preceding model for an explicit geometric…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Nanjie Yao , Gangjian Zhang , Wenhao Shen , Jian Shu , Hao Wang

Reconstructing interacting hands from monocular RGB data is a challenging task, as it involves many interfering factors, e.g. self- and mutual occlusion and similar textures. Previous works only leverage information from a single RGB image…

Computer Vision and Pattern Recognition · Computer Science 2024-01-08 Weichao Zhao , Hezhen Hu , Wengang Zhou , Li li , Houqiang Li

Reconstructing dynamic articulated objects from a singular monocular video is challenging, requiring joint estimation of shape, motion, and camera parameters from limited views. Current methods typically demand extensive computational…

Computer Vision and Pattern Recognition · Computer Science 2024-05-22 Hao Zhang , Fang Li , Samyak Rawlekar , Narendra Ahuja

We study the problem of imitating object interactions from Internet videos. This requires understanding the hand-object interactions in 4D, spatially in 3D and over time, which is challenging due to mutual hand-object occlusions. In this…

Computer Vision and Pattern Recognition · Computer Science 2022-11-24 Austin Patel , Andrew Wang , Ilija Radosavovic , Jitendra Malik