English

Hallucination as an Upper Bound: A New Perspective on Text-to-Image Evaluation

Computer Vision and Pattern Recognition 2025-11-11 v2 Computation and Language

Abstract

In language and vision-language models, hallucination is broadly understood as content generated from a model's prior knowledge or biases rather than from the given input. While this phenomenon has been studied in those domains, it has not been clearly framed for text-to-image (T2I) generative models. Existing evaluations mainly focus on alignment, checking whether prompt-specified elements appear, but overlook what the model generates beyond the prompt. We argue for defining hallucination in T2I as bias-driven deviations and propose a taxonomy with three categories: attribute, relation, and object hallucinations. This framing introduces an upper bound for evaluation and surfaces hidden biases, providing a foundation for richer assessment of T2I models.

Keywords

Cite

@article{arxiv.2509.21257,
  title  = {Hallucination as an Upper Bound: A New Perspective on Text-to-Image Evaluation},
  author = {Seyed Amir Kasaei and Mohammad Hossein Rohban},
  journal= {arXiv preprint arXiv:2509.21257},
  year   = {2025}
}

Comments

Accepted at GenProCC NeurIPS 2025 Workshop

R2 v1 2026-07-01T05:56:27.561Z