LLM 作为裁判器:有效性与可靠性的探讨
摘要
自然语言生成(NLG)系统的评估始终是自然语言处理(NLP)的核心挑战, further complicated by the rise of large language models (LLMs) that aims to be general-purpose. 最近,大型语言模型作为裁判器(LLJs) emerged as a promising alternative to traditional metrics, but their validity remains underexplored. This position paper argues that the current enthusiasm around LLJs may be premature, as their adoption has outpaced rigorous scrutiny of their reliability and validity as evaluators. Drawing on measurement theory from the social sciences, we identify and critically assess four core assumptions underlying the use of LLJs: their ability to act as proxies for human judgment, their capabilities as evaluators, their scalability, and their cost-effectiveness. We examine how each of these assumptions may be challenged by the inherent limitations of LLMs, LLJs, or current practices in NLG evaluation. To ground our analysis, we explore three applications of LLJs: text summarization, data annotation, and safety alignment. Finally, we highlight the need for more responsible evaluation practices in LLJs evaluation, to ensure that their growing role in the field supports, rather than undermines, progress in NLG.
引用
@article{arxiv.2508.18076,
title = {Neither Valid nor Reliable? Investigating the Use of LLMs as Judges},
author = {Khaoula Chehbouni and Mohammed Haddou and Jackie Chi Kit Cheung and Golnoosh Farnadi},
journal= {arXiv preprint arXiv:2508.18076},
year = {2025}
}
备注
Prepared for conference submission