English

Render-and-Compare: Cross-View 6 DoF Localization from Noisy Prior

Computer Vision and Pattern Recognition 2024-07-09 v2

Abstract

Despite the significant progress in 6-DoF visual localization, researchers are mostly driven by ground-level benchmarks. Compared with aerial oblique photography, ground-level map collection lacks scalability and complete coverage. In this work, we propose to go beyond the traditional ground-level setting and exploit the cross-view localization from aerial to ground. We solve this problem by formulating camera pose estimation as an iterative render-and-compare pipeline and enhancing the robustness through augmenting seeds from noisy initial priors. As no public dataset exists for the studied problem, we collect a new dataset that provides a variety of cross-view images from smartphones and drones and develop a semi-automatic system to acquire ground-truth poses for query images. We benchmark our method as well as several state-of-the-art baselines and demonstrate that our method outperforms other approaches by a large margin.

Keywords

Cite

@article{arxiv.2302.06287,
  title  = {Render-and-Compare: Cross-View 6 DoF Localization from Noisy Prior},
  author = {Shen Yan and Xiaoya Cheng and Yuxiang Liu and Juelin Zhu and Rouwan Wu and Yu Liu and Maojun Zhang},
  journal= {arXiv preprint arXiv:2302.06287},
  year   = {2024}
}
R2 v1 2026-06-28T08:38:39.726Z