English

To See or To Please: Uncovering Visual Sycophancy and Split Beliefs in VLMs

Computer Vision and Pattern Recognition 2026-05-27 v3 Artificial Intelligence

Abstract

When VLMs answer correctly, do they genuinely rely on visual information? We introduce a Tri-Layer Diagnostic Framework with three per-sample metrics: Latent Anomaly Detection, Visual Necessity Score, and Competition Score, which disentangle perception, dependency, and alignment failures. Across 9 VLMs and 9,000 model-sample pairs under counterfactual blind, noise, and conflict interventions, 72.9% of samples exhibit Visual Sycophancy, a Split Beliefs pattern in which internal evidence is preserved yet a hallucinated answer is decoded, while zero samples show Robust Refusal, indicating that current alignment training has eliminated refusal as a decoding outcome. Scaling within the Qwen-VL family, both within- and across-generation, monotonically reduces Language Shortcuts but amplifies Visual Sycophancy, showing that scale and newer post-training alone cannot resolve the grounding problem. Diagnostic scores further enable a training-free selective-prediction strategy yielding up to +9.5 percentage points accuracy at 50% coverage.

Keywords

Cite

@article{arxiv.2603.18373,
  title  = {To See or To Please: Uncovering Visual Sycophancy and Split Beliefs in VLMs},
  author = {Rui Hong and Shuxue Quan},
  journal= {arXiv preprint arXiv:2603.18373},
  year   = {2026}
}

Comments

14 pages, 1 figures