English

MV-TAP: Tracking Any Point in Multi-View Videos

Computer Vision and Pattern Recognition 2025-12-02 v1

Abstract

Multi-view camera systems enable rich observations of complex real-world scenes, and understanding dynamic objects in multi-view settings has become central to various applications. In this work, we present MV-TAP, a novel point tracker that tracks points across multi-view videos of dynamic scenes by leveraging cross-view information. MV-TAP utilizes camera geometry and a cross-view attention mechanism to aggregate spatio-temporal information across views, enabling more complete and reliable trajectory estimation in multi-view videos. To support this task, we construct a large-scale synthetic training dataset and real-world evaluation sets tailored for multi-view tracking. Extensive experiments demonstrate that MV-TAP outperforms existing point-tracking methods on challenging benchmarks, establishing an effective baseline for advancing research in multi-view point tracking.

Keywords

Cite

@article{arxiv.2512.02006,
  title  = {MV-TAP: Tracking Any Point in Multi-View Videos},
  author = {Jahyeok Koo and Inès Hyeonsu Kim and Mungyeom Kim and Junghyun Park and Seohyun Park and Jaeyeong Kim and Jung Yi and Seokju Cho and Seungryong Kim},
  journal= {arXiv preprint arXiv:2512.02006},
  year   = {2025}
}

Comments

Project Page: https://cvlab-kaist.github.io/MV-TAP/

R2 v1 2026-07-01T08:04:18.895Z