English

Point Clouds Are Specialized Images: A Knowledge Transfer Approach for 3D Understanding

Computer Vision and Pattern Recognition 2024-04-24 v2

Abstract

Self-supervised representation learning (SSRL) has gained increasing attention in point cloud understanding, in addressing the challenges posed by 3D data scarcity and high annotation costs. This paper presents PCExpert, a novel SSRL approach that reinterprets point clouds as "specialized images". This conceptual shift allows PCExpert to leverage knowledge derived from large-scale image modality in a more direct and deeper manner, via extensively sharing the parameters with a pre-trained image encoder in a multi-way Transformer architecture. The parameter sharing strategy, combined with a novel pretext task for pre-training, i.e., transformation estimation, empowers PCExpert to outperform the state of the arts in a variety of tasks, with a remarkable reduction in the number of trainable parameters. Notably, PCExpert's performance under LINEAR fine-tuning (e.g., yielding a 90.02% overall accuracy on ScanObjectNN) has already approached the results obtained with FULL model fine-tuning (92.66%), demonstrating its effective and robust representation capability.

Keywords

Cite

@article{arxiv.2307.15569,
  title  = {Point Clouds Are Specialized Images: A Knowledge Transfer Approach for 3D Understanding},
  author = {Jiachen Kang and Wenjing Jia and Xiangjian He and Kin Man Lam},
  journal= {arXiv preprint arXiv:2307.15569},
  year   = {2024}
}
R2 v1 2026-06-28T11:42:53.904Z