English

Bi-CoG: Bi-Consistency-Guided Self-Training for Vision-Language Models

Machine Learning 2026-05-26 v3

Abstract

Exploiting unlabeled data through semi-supervised learning (SSL) or leveraging pre-trained models via fine-tuning are two prevailing paradigms for addressing label-scarce scenarios. Recently, growing attention has been given to combining fine-tuning of pre-trained vision-language models (VLMs) with SSL, forming the emerging paradigm of semi-supervised fine-tuning. However, existing methods often suffer from model bias and hyperparameter sensitivity, due to reliance on prediction consistency or pre-defined confidence thresholds. To address these limitations, we propose a simple yet effective plug-and-play methodology named Bi-Co\underline{\textbf{Bi-Co}}nsistency-G\underline{\textbf{G}}uided Self-Training (Bi-CoG), which assigns high-quality and low-bias pseudo-labels, by simultaneously exploiting inter-model and intra-model consistency, along with an error-aware dynamic pseudo-label assignment strategy. Both theoretical analysis and extensive experiments over 14 datasets demonstrate the effectiveness of Bi-CoG, which consistently and significantly improves the performance of existing methods.

Keywords

Cite

@article{arxiv.2510.20477,
  title  = {Bi-CoG: Bi-Consistency-Guided Self-Training for Vision-Language Models},
  author = {Rui Zhu and Song-Lin Lv and Zi-Kang Wang and Lan-Zhe Guo},
  journal= {arXiv preprint arXiv:2510.20477},
  year   = {2026}
}

Comments

Accepted by IJCAI 2026

R2 v1 2026-07-01T07:01:58.699Z