English

Cross-View Feature Matching: Survey, Benchmarking, and Foundation-Model Perspectives

Machine Learning 2026-08-11 v1 Computer Vision and Pattern Recognition

Abstract

Cross-view feature matching aims to establish reliable correspondences across images with large viewpoint variations. Over the past decade, the field has evolved from task-specific models toward increasingly unified and generalizable correspondence models, with recent progress further driven by the emergence of vision foundation models (VFMs). Despite these advances, existing studies remain highly diverse in their problem formulations, model architectures, training paradigms, and evaluation protocols, making it difficult to obtain a unified understanding of the field. In this survey, we present a unified review of cross-view feature matching. We first introduce a structured taxonomy covering feature extraction, single-type feature matcher, multi-type feature matcher, VFMs based methods, training strategy and robust estimation, providing a coherent framework for analysis and comparison. We further examine recent advances, distilling key design principles and highlighting the shift toward unified and generalizable correspondence models. We also provide a unified experimental benchmarking of representative state-of-the-art methods under consistent protocols, enabling fair and comprehensive performance comparisons. In addition, we discuss open challenges and future directions, including efficiency, robustness under extreme conditions, and cross-domain generalization. This survey aims to provide a comprehensive and structured reference for understanding the evolution, current landscape, and future development of cross-view feature matching in the era of vision foundation models.

Keywords

Cite

@article{arxiv.2608.11093,
  title  = {Cross-View Feature Matching: Survey, Benchmarking, and Foundation-Model Perspectives},
  author = {Songlin Du and Xiaoyong Lu and Zeyu Wu and Xiaobo Lu and Guobao Xiao and Bin Fan and Jiayi Ma and Takeshi Ikenaga},
  journal= {arXiv preprint arXiv:2608.11093},
  year   = {2026}
}

Comments

This manuscript goes beyond a conventional survey. It proposes a new taxonomy for cross-view feature matching, provides extensive benchmarking under unified datasets and protocols, and offers original analysis from the perspective of vision foundation models. These contributions provide substantive methodological synthesis, empirical findings, and new research insights