English

Spotlight and Shadow: Attention-Guided Dual-Anchor Introspective Decoding for MLLM Hallucination Mitigation

Computer Vision and Pattern Recognition 2026-04-14 v1

Abstract

Multimodal Large Language Models (MLLMs) have demonstrated remarkable reasoning capabilities yet continue to suffer from hallucination, where generated text contradicts visual content. In this paper, we introduce Dual-Anchor Introspective Decoding (DaID), a novel contrastive decoding framework that dynamically calibrates each token generation by mining the model's internal perceptual discrepancies. Specifically, DaID identifies a Spotlight layer to amplify visual factual signals and a Shadow layer to suppress textual inertia. By leveraging visual attention distributions to guide this dual-anchor selection process, our method ensures precise, token-specific adaptation. Experimental results across multiple benchmarks and MLLMs demonstrate that DaID significantly mitigates hallucination while enhancing general reasoning capabilities.

Keywords

Cite

@article{arxiv.2604.10071,
  title  = {Spotlight and Shadow: Attention-Guided Dual-Anchor Introspective Decoding for MLLM Hallucination Mitigation},
  author = {Yebo Wu and Han Jin and Zhijiang Guo and Li Li},
  journal= {arXiv preprint arXiv:2604.10071},
  year   = {2026}
}

Comments

Accepted for Findings of ACL 2026

R2 v1 2026-07-01T12:04:08.573Z