English

SOTA: Self-adaptive Optimal Transport for Zero-Shot Classification with Multiple Foundation Models

Computer Vision and Pattern Recognition 2026-03-13 v3

Abstract

Foundation models have attracted widespread attention across domains due to their powerful zero-shot classification capabilities. This work is motivated by two key observations: (1) \textit{Vision-Language Models} (VLMs), such as CLIP, often over-rely on class-level textual priors and struggle to capture fine-grained visual cues, whereas \textit{Vision-only Foundation Models} (VFMs), such as DINO, provide rich and discriminative visual features but lack semantic alignment; (2) the performance of different VLMs varies considerably across datasets owing to differences in pre-training. To address these challenges, we propose \textbf{SOTA} (\textit{Self-adaptive Optimal TrAnsport}), a \textit{training-free} ensemble framework that integrates the outputs of multiple foundation models~(VFMs or VLMs) by learning a self-adaptive transport plan. Notably, \textbf{SOTA} is prior-free and automatically balances model contributions. Extensive experiments across diverse domains, including natural images, medical pathology, and remote sensing, validate the generalizability of \textbf{SOTA}. The results consistently show that it effectively leverages the complementary strengths of different foundation models and achieves substantial improvements over individual models. The implementation code is available at: https://github.com/Afleve/self-adaptive-Optimal-Transport.

Keywords

Cite

@article{arxiv.2506.13723,
  title  = {SOTA: Self-adaptive Optimal Transport for Zero-Shot Classification with Multiple Foundation Models},
  author = {Zhanxuan Hu and Qiyu Xu and Yu Duan and Yonghang Tai and Huafeng Li},
  journal= {arXiv preprint arXiv:2506.13723},
  year   = {2026}
}
R2 v1 2026-07-01T03:20:09.529Z