English

Improving the Robustness of 3D Human Pose Estimation: A Benchmark and Learning from Noisy Input

Computer Vision and Pattern Recognition 2024-04-17 v2

Abstract

Despite the promising performance of current 3D human pose estimation techniques, understanding and enhancing their generalization on challenging in-the-wild videos remain an open problem. In this work, we focus on the robustness of 2D-to-3D pose lifters. To this end, we develop two benchmark datasets, namely Human3.6M-C and HumanEva-I-C, to examine the robustness of video-based 3D pose lifters to a wide range of common video corruptions including temporary occlusion, motion blur, and pixel-level noise. We observe the poor generalization of state-of-the-art 3D pose lifters in the presence of corruption and establish two techniques to tackle this issue. First, we introduce Temporal Additive Gaussian Noise (TAGN) as a simple yet effective 2D input pose data augmentation. Additionally, to incorporate the confidence scores output by the 2D pose detectors, we design a confidence-aware convolution (CA-Conv) block. Extensively tested on corrupted videos, the proposed strategies consistently boost the robustness of 3D pose lifters and serve as new baselines for future research.

Keywords

Cite

@article{arxiv.2312.06797,
  title  = {Improving the Robustness of 3D Human Pose Estimation: A Benchmark and Learning from Noisy Input},
  author = {Trung-Hieu Hoang and Mona Zehni and Huy Phan and Duc Minh Vo and Minh N. Do},
  journal= {arXiv preprint arXiv:2312.06797},
  year   = {2024}
}
R2 v1 2026-06-28T13:47:42.479Z