English

FIHA: Autonomous Hallucination Evaluation in Vision-Language Models with Davidson Scene Graphs

Computer Vision and Pattern Recognition 2025-06-04 v2

Abstract

The rapid development of Large Vision-Language Models (LVLMs) often comes with widespread hallucination issues, making cost-effective and comprehensive assessments increasingly vital. Current approaches mainly rely on costly annotations and are not comprehensive -- in terms of evaluating all aspects such as relations, attributes, and dependencies between aspects. Therefore, we introduce the FIHA (autonomous Fine-graIned Hallucination evAluation evaluation in LVLMs), which could access hallucination LVLMs in the LLM-free and annotation-free way and model the dependency between different types of hallucinations. FIHA can generate Q&A pairs on any image dataset at minimal cost, enabling hallucination assessment from both image and caption. Based on this approach, we introduce a benchmark called FIHA-v1, which consists of diverse questions on various images from MSCOCO and Foggy. Furthermore, we use the Davidson Scene Graph (DSG) to organize the structure among Q&A pairs, in which we can increase the reliability of the evaluation. We evaluate representative models using FIHA-v1, highlighting their limitations and challenges. We released our code and data.

Keywords

Cite

@article{arxiv.2409.13612,
  title  = {FIHA: Autonomous Hallucination Evaluation in Vision-Language Models with Davidson Scene Graphs},
  author = {Bowen Yan and Zhengsong Zhang and Liqiang Jing and Eftekhar Hossain and Xinya Du},
  journal= {arXiv preprint arXiv:2409.13612},
  year   = {2025}
}

Comments

Accepted by Findings of ACL 2025

R2 v1 2026-06-28T18:51:34.186Z