中文

通过一致性训练减少政治操纵

计算与语言 2026-05-29 v2 人工智能

摘要

大型语言模型(LLM)在多种敏感情境下表现出系统性政治偏见。我们发现,LLM在处理来自对立政治立场的话题时表现出不对称性。我们称之为隐性政治偏见,并识别了其运作的7类技术。我们提出了两种隐性偏见度量指标:情感一致性衡量对称的修辞和叙事;帮助fulness一致性衡量对称的深度和参与度。为减少这两种隐性偏见,我们引入了政治一致性训练(PCT),这是一种RL训练方法,包含两种互补范式:情感一致性训练和帮助fulness一致性训练。我们表明,PCT在保持整体帮助fulness的同时,显著减少了隐性政治偏见,并在 held-out基准上实现泛化。我们在 https://political-manipulation.ai 公开了这项工作。

关键词

引用

@article{arxiv.2605.22771,
  title  = {Reducing Political Manipulation with Consistency Training},
  author = {Long Phan and Devin Kim and Alexander Pan and Alice Blair and Adam Khoja and Dan Hendrycks},
  journal= {arXiv preprint arXiv:2605.22771},
  year   = {2026}
}