English
Related papers

Related papers: Wild2Avatar: Rendering Humans Behind Occlusions

200 papers

Visual understanding of the world goes beyond the semantics and flat structure of individual images. In this work, we aim to capture both the 3D structure and dynamics of real-world scenes from monocular real-world videos. Our Dynamic Scene…

Computer Vision and Pattern Recognition · Computer Science 2024-03-18 Maximilian Seitzer , Sjoerd van Steenkiste , Thomas Kipf , Klaus Greff , Mehdi S. M. Sajjadi

Traditionally, video conferencing is a widely adopted solution for telecommunication, but a lack of immersiveness comes inherently due to the 2D nature of facial representation. The integration of Virtual Reality (VR) in a…

Computer Vision and Pattern Recognition · Computer Science 2021-12-03 Surabhi Gupta , Ashwath Shetty , Avinash Sharma

Existing augmented reality (AR) applications often ignore occlusion between real hands and virtual objects when incorporating virtual objects in our views. The challenges come from the lack of accurate depth and mismatch between real and…

Graphics · Computer Science 2020-06-24 Xiao Tang , Xiaowei Hu , Chi-Wing Fu , Daniel Cohen-Or

One major challenge for monocular 3D human pose estimation in-the-wild is the acquisition of training data that contains unconstrained images annotated with accurate 3D poses. In this paper, we address this challenge by proposing a…

Computer Vision and Pattern Recognition · Computer Science 2020-03-18 Umar Iqbal , Pavlo Molchanov , Jan Kautz

This paper introduces an unsupervised framework to extract semantically rich features for video representation. Inspired by how the human visual system groups objects based on motion cues, we propose a deep convolutional neural network that…

Computer Vision and Pattern Recognition · Computer Science 2017-07-18 Xunyu Lin , Victor Campos , Xavier Giro-i-Nieto , Jordi Torres , Cristian Canton Ferrer

Advances in Deep Learning have recently made it possible to recover full 3D meshes of human poses from individual images. However, extension of this notion to videos for recovering temporally coherent poses still remains unexplored. A major…

Computer Vision and Pattern Recognition · Computer Science 2019-07-02 Jian Liu , Naveed Akhtar , Ajmal Mian

Human motion recovery for real-world interaction demands both precise action details and metric-scale trajectories. Recovering absolute human pose from monocular input presents a viable solution, but faces two main challenges: (1) models'…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Zhumei Wang , Zechen Hu , Ruoxi Guo , Huaijin Pi , Ziyong Feng , Liang Zhang , Mingtao Pei , Siyuan Huang

We present the first approach to volumetric performance capture and novel-view rendering at real-time speed from monocular video, eliminating the need for expensive multi-view systems or cumbersome pre-acquisition of a personalized template…

Computer Vision and Pattern Recognition · Computer Science 2020-07-29 Ruilong Li , Yuliang Xiu , Shunsuke Saito , Zeng Huang , Kyle Olszewski , Hao Li

High-fidelity rendering of dynamic humans from monocular videos typically degrades catastrophically under occlusions. Existing solutions incorporate external priors-either hallucinating missing content via generative models, which induces…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Weiquan Wang , Feifei Shao , Lin Li , Zhen Wang , Jun Xiao , Long Chen

Egocentric "walking tour" videos provide a rich source of image data to develop rich and diverse visual models of environments around the world. However, the significant presence of humans in frames of these videos due to crowds and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Yujin Ham , Junho Kim , Vivek Boominathan , Guha Balakrishnan

Given a single RGB image of a complex outdoor road scene in the perspective view, we address the novel problem of estimating an occlusion-reasoned semantic scene layout in the top-view. This challenging problem not only requires an accurate…

Computer Vision and Pattern Recognition · Computer Science 2018-03-30 Samuel Schulter , Menghua Zhai , Nathan Jacobs , Manmohan Chandraker

The lack of occlusion data in common action recognition video datasets limits model robustness and hinders consistent performance gains. We build OccludeNet, a large-scale occluded video dataset including both real and synthetic occlusion…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Guanyu Zhou , Wenxuan Liu , Wenxin Huang , Xuemei Jia , Xian Zhong , Chia-Wen Lin

Tracking 3D human motion from egocentric multi-camera headset is challenged by severe egomotion, partial visibility or occlusions and lack of training data. Existing methods designed for monocular video often require static or slowly-moving…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Nan Yang , Julian Straub , Fan Zhang , Richard Newcombe , Jakob Engel , Lingni Ma

Holistic 3D human-scene reconstruction is a crucial and emerging research area in robot perception. A key challenge in holistic 3D human-scene reconstruction is to generate a physically plausible 3D scene from a single monocular RGB image.…

Computer Vision and Pattern Recognition · Computer Science 2023-07-28 Sandika Biswas , Kejie Li , Biplab Banerjee , Subhasis Chaudhuri , Hamid Rezatofighi

This paper presents a computational model to recover the most likely interpretation of the 3D scene structure from a planar image, where some objects may occlude others. The estimated scene interpretation is obtained by integrating some…

Computer Vision and Pattern Recognition · Computer Science 2016-03-30 Maria Oliver , Gloria Haro , Mariella Dimiccoli , Baptiste Mazin , Coloma Ballester

High-fidelity digital human representations are increasingly in demand in the digital world, particularly for interactive telepresence, AR/VR, 3D graphics, and the rapidly evolving metaverse. Even though they work well in small spaces,…

Computer Vision and Pattern Recognition · Computer Science 2024-10-24 Zexu Huang , Sarah Monazam Erfani , Siying Lu , Mingming Gong

3D human pose estimation (HPE) is crucial in many fields, such as human behavior analysis, augmented reality/virtual reality (AR/VR) applications, and self-driving industry. Videos that contain multiple potentially occluded people captured…

Computer Vision and Pattern Recognition · Computer Science 2020-11-03 Renshu Gu , Gaoang Wang , Jenq-Neng Hwang

To overcome the problem of occlusion in visual tracking, this paper proposes an occlusion-aware tracking algorithm. The proposed algorithm divides the object into discrete image patches according to the pixel distribution of the object by…

Computer Vision and Pattern Recognition · Computer Science 2021-04-19 Rongtai Caiand Peng Zhu

Reconstructing clothed humans from a single image is a fundamental task in computer vision with wide-ranging applications. Although existing monocular clothed human reconstruction solutions have shown promising results, they often rely on…

Computer Vision and Pattern Recognition · Computer Science 2025-10-20 Arindam Dutta , Meng Zheng , Zhongpai Gao , Benjamin Planche , Anwesha Choudhuri , Terrence Chen , Amit K. Roy-Chowdhury , Ziyan Wu

Robust 3D human pose estimation is crucial to ensure safe and effective human-robot collaboration. Accurate human perception,however, is particularly challenging in these scenarios due to strong occlusions and limited camera viewpoints.…

Computer Vision and Pattern Recognition · Computer Science 2024-08-29 Laura Bragagnolo , Matteo Terreran , Davide Allegro , Stefano Ghidoni