English

Quantifying Genuine Awareness in Hallucination Prediction Beyond Question-Side Shortcuts

Computation and Language 2026-03-11 v2

Abstract

Many works have proposed methodologies for language model (LM) hallucination detection and reported seemingly strong performance. However, we argue that the reported performance to date reflects not only a model's genuine awareness of its internal information, but also awareness derived purely from question-side information (e.g., benchmark hacking). While benchmark hacking can be effective for boosting hallucination detection score on existing benchmarks, it does not generalize to out-of-domain settings and practical usage. Nevertheless, disentangling how much of a model's hallucination detection performance arises from question-side awareness is non-trivial. To address this, we propose a methodology for measuring this effect without requiring human labor, Approximate Question-side Effect (AQE). Our analysis using AQE reveals that existing hallucination detection methods rely heavily on benchmark hacking.

Keywords

Cite

@article{arxiv.2509.15339,
  title  = {Quantifying Genuine Awareness in Hallucination Prediction Beyond Question-Side Shortcuts},
  author = {Yeongbin Seo and Dongha Lee and Jinyoung Yeo},
  journal= {arXiv preprint arXiv:2509.15339},
  year   = {2026}
}
R2 v1 2026-07-01T05:44:41.077Z