English

Hallucination Benchmark in Medical Visual Question Answering

Computation and Language 2024-04-04 v2 Artificial Intelligence Computer Vision and Pattern Recognition

Abstract

The recent success of large language and vision models (LLVMs) on vision question answering (VQA), particularly their applications in medicine (Med-VQA), has shown a great potential of realizing effective visual assistants for healthcare. However, these models are not extensively tested on the hallucination phenomenon in clinical settings. Here, we created a hallucination benchmark of medical images paired with question-answer sets and conducted a comprehensive evaluation of the state-of-the-art models. The study provides an in-depth analysis of current models' limitations and reveals the effectiveness of various prompting strategies.

Keywords

Cite

@article{arxiv.2401.05827,
  title  = {Hallucination Benchmark in Medical Visual Question Answering},
  author = {Jinge Wu and Yunsoo Kim and Honghan Wu},
  journal= {arXiv preprint arXiv:2401.05827},
  year   = {2024}
}

Comments

Accepted to ICLR 2024 Tiny Papers(Notable)

R2 v1 2026-06-28T14:14:09.738Z