English

Dynamic Camera Poses and Where to Find Them

Computer Vision and Pattern Recognition 2025-04-25 v1

Abstract

Annotating camera poses on dynamic Internet videos at scale is critical for advancing fields like realistic video generation and simulation. However, collecting such a dataset is difficult, as most Internet videos are unsuitable for pose estimation. Furthermore, annotating dynamic Internet videos present significant challenges even for state-of-theart methods. In this paper, we introduce DynPose-100K, a large-scale dataset of dynamic Internet videos annotated with camera poses. Our collection pipeline addresses filtering using a carefully combined set of task-specific and generalist models. For pose estimation, we combine the latest techniques of point tracking, dynamic masking, and structure-from-motion to achieve improvements over the state-of-the-art approaches. Our analysis and experiments demonstrate that DynPose-100K is both large-scale and diverse across several key attributes, opening up avenues for advancements in various downstream applications.

Keywords

Cite

@article{arxiv.2504.17788,
  title  = {Dynamic Camera Poses and Where to Find Them},
  author = {Chris Rockwell and Joseph Tung and Tsung-Yi Lin and Ming-Yu Liu and David F. Fouhey and Chen-Hsuan Lin},
  journal= {arXiv preprint arXiv:2504.17788},
  year   = {2025}
}

Comments

Accepted to CVPR 2025. Project Page: https://research.nvidia.com/labs/dir/dynpose-100k

R2 v1 2026-06-28T23:10:22.720Z