中文

从委托人到受托人:通过优化长期利益塑造LLM中的偏见与对齐

计算机与社会 2025-11-18 v2 人工智能

摘要

大型语言模型(LLM)在预测调查响应和政策偏好方面显示出前景广阔,这增加了他们在各个领域代表人类利益的潜力。目前大多数研究都专注于"行为克隆”,即评估模型在多大程度上复制了个人表达的偏好。drawing on theories of political representation, we highlight an underexplored design trade-off: whether AI systems should act as delegates, mirroring expressed preferences, or as trustees, exercising judgment about what best serves an individual's interests. This trade-off is closely related to issues of LLM sycophancy, where models can encourage behavior or validate beliefs that may be aligned with a user's short-term preferences, but is detrimental to their long-term interests. Through a series of experiments simulating votes on various policy issues in the U.S. context, we apply a temporal utility framework that weighs short and long-term interests (simulating a trustee role) and compare voting outcomes to behavior-cloning models (simulating a delegate). We find that trustee-style predictions weighted toward long-term interests produce policy decisions that align more closely with expert consensus on well-understood issues, but also show greater bias toward models' default stances on topics lacking clear agreement. These findings reveal a fundamental trade-off in designing AI systems to represent human interests. Delegate models better preserve user autonomy but may diverge from well-supported policy positions, while trustee models can promote welfare on well-understood issues yet risk paternalism and bias on subjective topics.

关键词

引用

@article{arxiv.2510.12689,
  title  = {From Delegates to Trustees: How Optimizing for Long-Term Interests Shapes Bias and Alignment in LLM},
  author = {Suyash Fulay and Jocelyn Zhu and Michiel Bakker},
  journal= {arXiv preprint arXiv:2510.12689},
  year   = {2025}
}