我认为自己不够格?——面向 LLM 招聘评估的语言习惯检测基准
计算与语言
2025-08-08 v1
摘要
本文引入了一个全面的基准,用于评估大型语言模型 (LLM) 对语言习惯的响应:这些是能够不经意地揭示性别、社会阶层或地区背景等人口属性的细微语言标记。通过使用 100 对经验证的提问-回答对,我们展示了 LLM 如何系统性地对特定语言模式进行惩罚,尤其是缓和语言,尽管内容质量相当。我们的基准生成受控的语言变体,以隔离特定现象 while maintaining semantic equivalence, which enables the precise measurement of demographic bias in automated evaluation systems. We validate our approach along multiple linguistic dimensions, showing that hedged responses receive 25.6% lower ratings on average, and demonstrate the benchmark's effectiveness in identifying model-specific biases. This work establishes a foundational framework for detecting and measuring linguistic discrimination in AI systems, with broad applications to fairness in automated decision-making contexts.
引用
@article{arxiv.2508.04939,
title = {I Think, Therefore I Am Under-Qualified? A Benchmark for Evaluating Linguistic Shibboleth Detection in LLM Hiring Evaluations},
author = {Julia Kharchenko and Tanya Roosta and Aman Chadha and Chirag Shah},
journal= {arXiv preprint arXiv:2508.04939},
year = {2025}
}