English
Related papers

Related papers: Monocular Real-Time Volumetric Performance Capture

200 papers

High-fidelity face digitization solutions often combine multi-view stereo (MVS) techniques for 3D reconstruction and a non-rigid registration step to establish dense correspondence across identities and expressions. A common problem is the…

Computer Vision and Pattern Recognition · Computer Science 2021-10-07 Tianye Li , Shichen Liu , Timo Bolkart , Jiayi Liu , Hao Li , Yajie Zhao

We propose an approach for reconstructing free-moving object from a monocular RGB video. Most existing methods either assume scene prior, hand pose prior, object category pose prior, or rely on local optimization with multiple sequence…

Computer Vision and Pattern Recognition · Computer Science 2024-05-13 Haixin Shi , Yinlin Hu , Daniel Koguciuk , Juan-Ting Lin , Mathieu Salzmann , David Ferstl

We propose a method to learn a high-quality implicit 3D head avatar from a monocular RGB video captured in the wild. The learnt avatar is driven by a parametric face model to achieve user-controlled facial expressions and head poses. Our…

Single-image 3D human reconstruction aims to reconstruct the 3D textured surface of the human body given a single image. While implicit function-based methods recently achieved reasonable reconstruction performance, they still bear…

Computer Vision and Pattern Recognition · Computer Science 2022-11-10 Gyumin Shim , Minsoo Lee , Jaegul Choo

This paper presents an approach that reconstructs a hand-held object from a monocular video. In contrast to many recent methods that directly predict object geometry by a trained network, the proposed approach does not require any learned…

Computer Vision and Pattern Recognition · Computer Science 2022-12-01 Di Huang , Xiaopeng Ji , Xingyi He , Jiaming Sun , Tong He , Qing Shuai , Wanli Ouyang , Xiaowei Zhou

Reconstructing 4D dynamic scenes from casually captured monocular videos is valuable but highly challenging, as each timestamp is observed from a single viewpoint. We introduce Vivid4D, a novel approach that enhances 4D monocular video…

Computer Vision and Pattern Recognition · Computer Science 2025-04-21 Jiaxin Huang , Sheng Miao , BangBang Yang , Yuewen Ma , Yiyi Liao

A monocular 3D object tracking system generally has only up-to-scale pose estimation results without any prior knowledge of the tracked object. In this paper, we propose a novel idea to recover the metric scale of an arbitrary dynamic…

Robotics · Computer Science 2018-08-22 Kejie Qiu , Tong Qin , Hongwen Xie , Shaojie Shen

Estimating a scene's depth to achieve collision avoidance against moving pedestrians is a crucial and fundamental problem in the robotic field. This paper proposes a novel, low complexity network architecture for fast and accurate human…

Computer Vision and Pattern Recognition · Computer Science 2021-08-25 Shan An , Fangru Zhou , Mei Yang , Haogang Zhu , Changhong Fu , Konstantinos A. Tsintotas

We present V2M4, a novel 4D reconstruction method that directly generates a usable 4D mesh animation asset from a single monocular video. Unlike existing approaches that rely on priors from multi-view image and video generation models, our…

Graphics · Computer Science 2025-07-30 Jianqi Chen , Biao Zhang , Xiangjun Tang , Peter Wonka

Self-supervised monocular depth estimation has gathered notable interest since it can liberate training from dependency on depth annotations. In monocular video training case, recent methods only conduct view synthesis between existing…

Computer Vision and Pattern Recognition · Computer Science 2024-07-22 Jinfeng Liu , Lingtong Kong , Bo Li , Zerong Wang , Hong Gu , Jinwei Chen

Recovering 3D full-body human pose is a challenging problem with many applications. It has been successfully addressed by motion capture systems with body worn markers and multiple cameras. In this paper, we address the more challenging…

Computer Vision and Pattern Recognition · Computer Science 2018-03-12 Xiaowei Zhou , Menglong Zhu , Georgios Pavlakos , Spyridon Leonardos , Kostantinos G. Derpanis , Kostas Daniilidis

Recovering 4D human-object interaction (HOI) from monocular video is a key step toward scalable 3D content creation, embodied AI, and simulation-based learning. Recent methods can reconstruct temporally coherent human and object…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Yubo Zhao , Yujin Chai , Yunao Dong , Chengfeng Zhao , Zijiao Zeng , Yuan Liu , Chi-Keung Tang

In this paper we present a novel method to estimate 3D human pose and shape from monocular videos. This task requires directly recovering pixel-alignment 3D human pose and body shape from monocular images or videos, which is challenging due…

Computer Vision and Pattern Recognition · Computer Science 2023-03-02 Sen Yang , Wen Heng , Gang Liu , Guozhong Luo , Wankou Yang , Gang Yu

We propose a novel neural rendering pipeline, Hybrid Volumetric-Textural Rendering (HVTR), which synthesizes virtual human avatars from arbitrary poses efficiently and at high quality. First, we learn to encode articulated human motions on…

Computer Vision and Pattern Recognition · Computer Science 2022-09-02 Tao Hu , Tao Yu , Zerong Zheng , He Zhang , Yebin Liu , Matthias Zwicker

This paper proposes a method to reconstruct the neural radiance field with equirectangular omnidirectional images. Implicit neural scene representation with a radiance field can reconstruct the 3D shape of a scene continuously within a…

Computer Vision and Pattern Recognition · Computer Science 2022-08-15 Qiaoge Li , Itsuki Ueda , Chun Xie , Hidehiko Shishido , Itaru Kitahara

Monocular dynamic video reconstruction faces significant challenges in dynamic human scenes due to geometric inconsistencies and resolution degradation issues. Existing methods lack 3D human structural understanding, producing geometrically…

Computer Vision and Pattern Recognition · Computer Science 2025-12-10 Weitao Xiong , Zhiyuan Yuan , Jiahao Lu , Chengfeng Zhao , Peng Li , Yuan Liu

Real-time rendering and animation of humans is a core function in games, movies, and telepresence applications. Existing methods have a number of drawbacks we aim to address with our work. Triangle meshes have difficulty modeling thin…

We introduce D$^3$-Human, a method for reconstructing Dynamic Disentangled Digital Human geometry from monocular videos. Past monocular video human reconstruction primarily focuses on reconstructing undecoupled clothed human bodies or only…

Computer Vision and Pattern Recognition · Computer Science 2025-01-06 Honghu Chen , Bo Peng , Yunfan Tao , Juyong Zhang

Motivated by augmented and virtual reality applications such as telepresence, there has been a recent focus in real-time performance capture of humans under motion. However, given the real-time constraint, these systems often suffer from…

We propose a method for in-hand 3D scanning of an unknown object with a monocular camera. Our method relies on a neural implicit surface representation that captures both the geometry and the appearance of the object, however, by contrast…

Computer Vision and Pattern Recognition · Computer Science 2023-06-23 Shreyas Hampali , Tomas Hodan , Luan Tran , Lingni Ma , Cem Keskin , Vincent Lepetit