Gaming the Answer Matcher: Examining the Impact of Text Manipulation on Automated Judgment
计算与语言
2026-01-15 v1
摘要
自动化答案匹配技术利用LLMs评估自由文本响应,通过将其与参考答案进行比较,显示出作为可扩展且对齐的替代方案的巨大潜力。然而,其可靠性需要对鲁棒性进行评估,以应对诸如猜测或冗长化等策略性攻击,这些攻击可能人为地提高评分而不提高实际正确性。在本文中,我们系统性地调查了此类策略是否欺骗答案匹配模型,通过引导 examinee模型:(1)生成冗长响应,(2)在不确定时提供多个答案,(3)在响应开头嵌入冲突答案与正确答案。我们的结果表明,这些操纵不会增加评分,甚至常常会降低评分。此外,二元评分(要求匹配器以明确的"正确"或"错误"作答)比连续评分(要求匹配器确定部分正确性)更抗攻击。这些发现表明,答案匹配通常对廉价文本操纵具有较强鲁棒性,在参考答案可用时,作为传统LLM-as-a-judge或人类评估的可行替代方案。
引用
@article{arxiv.2601.08849,
title = {Gaming the Answer Matcher: Examining the Impact of Text Manipulation on Automated Judgment},
author = {Manas Khatore and Sumana Sridharan and Kevork Sulahian and Benjamin J. Smith and Shi Feng},
journal= {arXiv preprint arXiv:2601.08849},
year = {2026}
}
备注
Accepted to the AAAI 2026 Workshop on AI Governance (AIGOV)