English

DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models

Computer Vision and Pattern Recognition 2025-08-01 v2 Artificial Intelligence

Abstract

Large vision-language models (LVLMs) have demonstrated exceptional performance on complex multimodal tasks. However, they continue to suffer from significant hallucination issues, including object, attribute, and relational hallucinations. To accurately detect these hallucinations, we investigated the variations in cross-modal attention patterns between hallucination and non-hallucination states. Leveraging these distinctions, we developed a lightweight detector capable of identifying hallucinations. Our proposed method, Detecting Hallucinations by Cross-modal Attention Patterns (DHCP), is straightforward and does not require additional LVLM training or extra LVLM inference steps. Experimental results show that DHCP achieves remarkable performance in hallucination detection. By offering novel insights into the identification and analysis of hallucinations in LVLMs, DHCP contributes to advancing the reliability and trustworthiness of these models. The code is available at https://github.com/btzyd/DHCP.

Keywords

Cite

@article{arxiv.2411.18659,
  title  = {DHCP: Detecting Hallucinations by Cross-modal Attention Pattern in Large Vision-Language Models},
  author = {Yudong Zhang and Ruobing Xie and Xingwu Sun and Yiqing Huang and Jiansheng Chen and Zhanhui Kang and Di Wang and Yu Wang},
  journal= {arXiv preprint arXiv:2411.18659},
  year   = {2025}
}

Comments

Accepted by ACM Multimedia 2025

R2 v1 2026-06-28T20:15:05.894Z