English

D2-Net: A Trainable CNN for Joint Detection and Description of Local Features

Computer Vision and Pattern Recognition 2019-05-10 v1

Abstract

In this work we address the problem of finding reliable pixel-level correspondences under difficult imaging conditions. We propose an approach where a single convolutional neural network plays a dual role: It is simultaneously a dense feature descriptor and a feature detector. By postponing the detection to a later stage, the obtained keypoints are more stable than their traditional counterparts based on early detection of low-level structures. We show that this model can be trained using pixel correspondences extracted from readily available large-scale SfM reconstructions, without any further annotations. The proposed method obtains state-of-the-art performance on both the difficult Aachen Day-Night localization dataset and the InLoc indoor localization benchmark, as well as competitive performance on other benchmarks for image matching and 3D reconstruction.

Keywords

Cite

@article{arxiv.1905.03561,
  title  = {D2-Net: A Trainable CNN for Joint Detection and Description of Local Features},
  author = {Mihai Dusmanu and Ignacio Rocco and Tomas Pajdla and Marc Pollefeys and Josef Sivic and Akihiko Torii and Torsten Sattler},
  journal= {arXiv preprint arXiv:1905.03561},
  year   = {2019}
}

Comments

Accepted at CVPR 2019

R2 v1 2026-06-23T09:01:36.590Z