Scaling Vision-Language Models Is Not Enough to Mitigate Bias
Abstract
Vision-Language Models (VLMs) such as CLIP are now foundational to multimodal systems, yet their robustness to spurious correlations remains poorly understood at scale. We present the first large-scale empirical study of 194 publicly available VLMs, including 16 model families, covering a wide range of model sizes, 24 training datasets, and three evaluation benchmarks, namely ImageNet (overall performance), CelebA (typical single-attribute bias), and UrbanCars (complex multi-attribute biases). Across these settings, the Spearman correlation between model scale and performance weakens as evaluation shifts from ImageNet () to single-attribute () and further to multi-attribute () bias benchmarks. In contrast, properties of the training data (size and quality) show more consistent relationships with worst-group accuracy across both bias benchmarks. Notably, curated datasets yield improvements of up to 25% over uncurated alternatives at a comparable scale. Finally, the effect of architectural choices (e.g., patch size, image resolution) is highly context-dependent, varying with the nature of the benchmark, including the type of bias and its spatial distribution within images.
Cite
@article{arxiv.2607.28211,
title = {Scaling Vision-Language Models Is Not Enough to Mitigate Bias},
author = {Ioannis Sarridis and Ioannis Kompatsiaris and Symeon Papadopoulos},
journal= {arXiv preprint arXiv:2607.28211},
year = {2026}
}