评估与缓解 LLM-as-a-judge 在通信系统中的偏差
人工智能
2026-03-03 v3 密码学与安全
摘要
大语言模型 (LLM) 越来越被用于自动评估通信系统中内容质量,例如评估 telecom customer support chatbot 中的响应质量。然而,AI "裁判" 的 impartiality 并非必然,任何其评估准则中的偏差都可能扭曲结果并削弱用户信任。在本文中,我们系统性地调查了 6 个 LLM-as-a-judge 模型(涵盖 prompt-based 和 fine-tuned judges)在 pointwise scoring setting 下的判断偏差,涵盖 11 种偏差类型,包括隐式和显式形式。我们观察到 state-of-the-art LLM 裁判对偏差输入表现出鲁健性,通常会为其分配低于相应 clean samples 的分数。我们进一步发现,judged 分数与 task difficulty 相关联:像 GPQA 这类困难数据集会yield 更低的平均分数,而开放式推理数据集(如 JudgeLM-val)则看到更高的平均分数。最后,我们提出了四种潜在的缓解策略,以确保在实际通信场景中实现公平可靠的 AI 裁判。
引用
@article{arxiv.2510.12462,
title = {Evaluating and Mitigating LLM-as-a-judge Bias in Communication Systems},
author = {Jiaxin Gao and Chen Chen and Yanwen Jia and Xueluan Gong and Kwok-Yan Lam and Qian Wang},
journal= {arXiv preprint arXiv:2510.12462},
year = {2026}
}