English

On the power of data augmentation for head pose estimation

Computer Vision and Pattern Recognition 2024-10-17 v3

Abstract

Deep learning has been impressively successful in the last decade in predicting human head poses from monocular images. However, for in-the-wild inputs the research community relies predominantly on a single training set, 300W-LP, of semisynthetic nature without many alternatives. This paper focuses on gradual extension and improvement of the data to explore the performance achievable with augmentation and synthesis strategies further. Modeling-wise a novel multitask head/loss design which includes uncertainty estimation is proposed. Overall, the thus obtained models are small, efficient, suitable for full 6 DoF pose estimation, and exhibit very competitive accuracy.

Keywords

Cite

@article{arxiv.2407.05357,
  title  = {On the power of data augmentation for head pose estimation},
  author = {Michael Welter},
  journal= {arXiv preprint arXiv:2407.05357},
  year   = {2024}
}

Comments

CVPR version. Added evaluation on BIWI. Plenty of writing changes

R2 v1 2026-06-28T17:31:52.886Z