English

Human Performance Capture from Monocular Video in the Wild

Computer Vision and Pattern Recognition 2021-12-01 v2

Abstract

Capturing the dynamically deforming 3D shape of clothed human is essential for numerous applications, including VR/AR, autonomous driving, and human-computer interaction. Existing methods either require a highly specialized capturing setup, such as expensive multi-view imaging systems, or they lack robustness to challenging body poses. In this work, we propose a method capable of capturing the dynamic 3D human shape from a monocular video featuring challenging body poses, without any additional input. We first build a 3D template human model of the subject based on a learned regression model. We then track this template model's deformation under challenging body articulations based on 2D image observations. Our method outperforms state-of-the-art methods on an in-the-wild human video dataset 3DPW. Moreover, we demonstrate its efficacy in robustness and generalizability on videos from iPER datasets.

Keywords

Cite

@article{arxiv.2111.14672,
  title  = {Human Performance Capture from Monocular Video in the Wild},
  author = {Chen Guo and Xu Chen and Jie Song and Otmar Hilliges},
  journal= {arXiv preprint arXiv:2111.14672},
  year   = {2021}
}
R2 v1 2026-06-24T07:56:00.364Z