基准是否低估了 LLM 性能?基于 LLM-首选人类裁决评估的幻觉检测
摘要
幻觉仍是大型语言模型 (LLM) 的持久挑战,特别是在上下文依赖的设置中,例如 RAG 和代理 AI 系统中。本研究聚焦于总结任务中的上下文幻觉检测。我们通过比较原始基准注释与来自 Gemini 2.5 Flash 和 GPT-5 Mini 的基于原因和片段的预测,分析了 QAGS-C 和 SummEval 数据集。为解决人类标签与 LLM 判断之间的系统性差异,我们通过包含 2 位跨文化裁决员的事人裁决过程重新评估了所有存在冲突的样本。ollowing this re-evaluation, triple agreement (between human, GPT, and Gemini) increased by 6.38% for QAGS-C and 7.62% for SummEval。 Similarly, model accuracy improved, with GPT increasing by 4.25% on QAGS-C and 2.34% on SummEval, while Gemini showed gains of 8.51% and 3.80%, respectively. Notably, adjudicators frequently sided with the models' judgments over original human annotations when LLMs provided explicit reasoning. Overall human adjudicator agreement ranged between 83% and 87%。These findings suggest that for ambiguity-prone tasks, single-pass annotations may be insufficient, and model-assisted re-evaluation yields more reliable benchmarks.
引用
@article{arxiv.2605.08462,
title = {Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment},
author = {I. F. Atasoy and B. Mutlu and E. A. Sezer and A. Wahdan},
journal= {arXiv preprint arXiv:2605.08462},
year = {2026}
}
备注
Presented at the ROMCIR Workshop at ECIR 2026