English

LCD: Learned Cross-Domain Descriptors for 2D-3D Matching

Computer Vision and Pattern Recognition 2019-11-22 v1

Abstract

In this work, we present a novel method to learn a local cross-domain descriptor for 2D image and 3D point cloud matching. Our proposed method is a dual auto-encoder neural network that maps 2D and 3D input into a shared latent space representation. We show that such local cross-domain descriptors in the shared embedding are more discriminative than those obtained from individual training in 2D and 3D domains. To facilitate the training process, we built a new dataset by collecting 1.4\approx 1.4 millions of 2D-3D correspondences with various lighting conditions and settings from publicly available RGB-D scenes. Our descriptor is evaluated in three main experiments: 2D-3D matching, cross-domain retrieval, and sparse-to-dense depth estimation. Experimental results confirm the robustness of our approach as well as its competitive performance not only in solving cross-domain tasks but also in being able to generalize to solve sole 2D and 3D tasks. Our dataset and code are released publicly at \url{https://hkust-vgd.github.io/lcd}.

Keywords

Cite

@article{arxiv.1911.09326,
  title  = {LCD: Learned Cross-Domain Descriptors for 2D-3D Matching},
  author = {Quang-Hieu Pham and Mikaela Angelina Uy and Binh-Son Hua and Duc Thanh Nguyen and Gemma Roig and Sai-Kit Yeung},
  journal= {arXiv preprint arXiv:1911.09326},
  year   = {2019}
}

Comments

Accepted to AAAI 2020 (Oral)

R2 v1 2026-06-23T12:23:05.104Z