中文

基准是否低估了 LLM 性能?基于 LLM-首选人类裁决评估的幻觉检测

计算与语言 2026-05-12 v1 人工智能

摘要

幻觉仍是大型语言模型 (LLM) 的持久挑战,特别是在上下文依赖的设置中,例如 RAG 和代理 AI 系统中。本研究聚焦于总结任务中的上下文幻觉检测。我们通过比较原始基准注释与来自 Gemini 2.5 Flash 和 GPT-5 Mini 的基于原因和片段的预测,分析了 QAGS-C 和 SummEval 数据集。为解决人类标签与 LLM 判断之间的系统性差异,我们通过包含 2 位跨文化裁决员的事人裁决过程重新评估了所有存在冲突的样本。ollowing this re-evaluation, triple agreement (between human, GPT, and Gemini) increased by 6.38% for QAGS-C and 7.62% for SummEval。 Similarly, model accuracy improved, with GPT increasing by 4.25% on QAGS-C and 2.34% on SummEval, while Gemini showed gains of 8.51% and 3.80%, respectively. Notably, adjudicators frequently sided with the models' judgments over original human annotations when LLMs provided explicit reasoning. Overall human adjudicator agreement ranged between 83% and 87%。These findings suggest that for ambiguity-prone tasks, single-pass annotations may be insufficient, and model-assisted re-evaluation yields more reliable benchmarks.

关键词

引用

@article{arxiv.2605.08462,
  title  = {Do Benchmarks Underestimate LLM Performance? Evaluating Hallucination Detection With LLM-First Human-Adjudicated Assessment},
  author = {I. F. Atasoy and B. Mutlu and E. A. Sezer and A. Wahdan},
  journal= {arXiv preprint arXiv:2605.08462},
  year   = {2026}
}

备注

Presented at the ROMCIR Workshop at ECIR 2026