English

Pic@Point: Cross-Modal Learning by Local and Global Point-Picture Correspondence

Computer Vision and Pattern Recognition 2024-10-15 v1 Artificial Intelligence

Abstract

Self-supervised pre-training has achieved remarkable success in NLP and 2D vision. However, these advances have yet to translate to 3D data. Techniques like masked reconstruction face inherent challenges on unstructured point clouds, while many contrastive learning tasks lack in complexity and informative value. In this paper, we present Pic@Point, an effective contrastive learning method based on structural 2D-3D correspondences. We leverage image cues rich in semantic and contextual knowledge to provide a guiding signal for point cloud representations at various abstraction levels. Our lightweight approach outperforms state-of-the-art pre-training methods on several 3D benchmarks.

Keywords

Cite

@article{arxiv.2410.09519,
  title  = {Pic@Point: Cross-Modal Learning by Local and Global Point-Picture Correspondence},
  author = {Vencia Herzog and Stefan Suwelack},
  journal= {arXiv preprint arXiv:2410.09519},
  year   = {2024}
}

Comments

Accepted at ACML 2024