English

System-Mediated Attention Imbalances Make Vision-Language Models Say Yes

Computation and Language 2026-04-27 v2

Abstract

Vision-language model (VLM) hallucination is commonly linked to imbalanced allocation of attention across input modalities: system, image and text. However, existing mitigation strategies tend towards an image-centric interpretation of these imbalances, often prioritising increased image attention while giving less consideration to the roles of the other modalities. In this study, we evaluate a more holistic, system-mediated account, which attributes these imbalances to functionally redundant system weights that reduce attention to image and textual inputs. We show that this framework offers a useful empirical perspective on the yes-bias, a common form of hallucination in which VLMs indiscriminately respond `yes'. Causally redistributing attention from the system modality to image and textual inputs substantially suppresses this bias, often outperforming existing approaches. We further present evidence suggesting that system-mediated attention imbalances contribute to the yes-bias by encouraging a default reliance on coarse input representations, which are effective for some tasks but ill-suited to others. Taken together, these findings firmly establish system attention as a key factor in VLM hallucination and highlight its potential as a lever for mitigation.

Keywords

Cite

@article{arxiv.2601.12430,
  title  = {System-Mediated Attention Imbalances Make Vision-Language Models Say Yes},
  author = {Tsan Tsai Chan and Varsha Suresh and Anisha Saha and Michael Hahn and Vera Demberg},
  journal= {arXiv preprint arXiv:2601.12430},
  year   = {2026}
}

Comments

Accepted to ACL Findings 2026

R2 v1 2026-07-01T09:09:32.770Z