English

Training a Student Expert via Semi-Supervised Foundation Model Distillation

Computer Vision and Pattern Recognition 2026-04-07 v1

Abstract

Foundation models deliver strong perception but are often too computationally heavy to deploy, and adapting them typically requires costly annotations. We introduce a semi-supervised knowledge distillation (SSKD) framework that compresses pre-trained vision foundation models (VFMs) into compact experts using limited labeled and abundant unlabeled data, and instantiate it for instance segmentation where per-pixel labels are particularly expensive. The framework unfolds in three stages: (1) domain adaptation of the VFM(s) via self-training with contrastive calibration, (2) knowledge transfer through a unified multi-objective loss, and (3) student refinement to mitigate residual pseudo-label bias. Central to our approach is an instance-aware pixel-wise contrastive loss that fuses mask and class scores to extract informative negatives and enforce clear inter-instance margins. By maintaining this contrastive signal across both adaptation and distillation, we align teacher and student embeddings and more effectively leverage unlabeled images. On Cityscapes and ADE20K, our 11×\approx 11\times smaller student improves over its zero-shot VFM teacher(s) by +11.9 and +8.6 AP, surpasses adapted teacher(s) by +3.4 and +1.5 AP, and outperforms state-of-the-art SSKD methods on benchmarks.

Keywords

Cite

@article{arxiv.2604.03841,
  title  = {Training a Student Expert via Semi-Supervised Foundation Model Distillation},
  author = {Pardis Taghavi and Tian Liu and Renjie Li and Reza Langari and Zhengzhong Tu},
  journal= {arXiv preprint arXiv:2604.03841},
  year   = {2026}
}

Comments

Accepted to the 2026 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). 14 pages, 9 figures

R2 v1 2026-07-01T11:54:03.295Z