心意相差:语言模型中陈述-显露偏好差距的影响因素
人工智能
2026-05-14 v2 新兴技术
摘要
近期研究识别了语言模型(LMs)中的陈述-显露(stated-revealed, SvR)偏好差距:即模型所倾导的值与其在情境下的选择之间的不匹配。现有的评估高度依赖二进制强制选择提示,这将真实偏好与 elicitation protocol 的副作用捆绑在一起。我们系统性地研究了 elicitation protocol 对 SvR 相关性的影响,涵盖 24 个 LMs。允许在陈述偏好 elicitation 中包含中立和放弃选项,可排除弱信号,显著提升 Spearman 秩相关系数()在自愿陈述偏好与强制选择显露偏好之间。然而,进一步允许在显露偏好 elicitation 中放弃选项,会导致 接近零或负值,由于高中立率。最后,我们发现使用陈述偏好引导 system prompt 在显露偏好 elicitation 中并不可靠地提高 SvR 相关性,尤其是在 AIRiskDilemmas 上。总体而言,我们的结果表明 SvR 相关性高度依赖于 protocol,且偏好 elicitation 需要方法能考虑不可确定偏好的事实。
引用
@article{arxiv.2601.21975,
title = {Mind the Gap: How Elicitation Protocols Shape the Stated-Revealed Preference Gap in Language Models},
author = {Pranav Mahajan and Ihor Kendiukhov and Syed Hussain and Lydia Nottingham},
journal= {arXiv preprint arXiv:2601.21975},
year = {2026}
}
备注
Accepted to ACL 2026 Eval Eval Workshop and 3rd Technical AI Safety Conference (TAIS 2026)