English

Fine-Grained Classification: Connecting Metadata via Cross-Contrastive Pre-Training

Computer Vision and Pattern Recognition 2025-10-21 v2 Machine Learning

Abstract

Fine-grained visual classification aims to recognize objects belonging to many subordinate categories of a supercategory, where appearance alone often fails to distinguish highly similar classes. We propose a unified framework that integrates image, text, and metadata via cross-contrastive pre-training. We first align the three modality encoders in a shared embedding space and then fine-tune the image and metadata encoders for classification. On NABirds, our approach improves over the baseline by 7.83% and achieves 84.44% top-1 accuracy, outperforming strong multimodal methods.

Keywords

Cite

@article{arxiv.2504.20322,
  title  = {Fine-Grained Classification: Connecting Metadata via Cross-Contrastive Pre-Training},
  author = {Sumit Mamtani and Yash Thesia},
  journal= {arXiv preprint arXiv:2504.20322},
  year   = {2025}
}

Comments

5 pages, 4 figures. Accepted at IEEE ISCMI 2025

R2 v1 2026-06-28T23:14:36.509Z