中文

价值对齐:一种形式化方法

人工智能 2024-02-08 v1

摘要

原则应当治理自主 AI 系统。它本质上指出,系统的目标和行为应与人类价值观对齐。但如何确保价值对齐?在本文中,我们首先提供了一个形式化模型,通过偏好来表示价值,并提供了计算价值聚合的方法;即关于智能体群体的偏好和/或关于价值集合的偏好。随后,针对给定的规范和给定的价值,通过其导致世界未来状态偏好的增减来定义并计算价值对齐。我们聚焦于规范,因为正是规范治理着行为,因此,给定系统与给定价值的对齐将由系统遵循的规范所决定。

关键词

引用

@article{arxiv.2110.09240,
  title  = {Value alignment: a formal approach},
  author = {Carles Sierra and Nardine Osman and Pablo Noriega and Jordi Sabater-Mir and Antoni Perelló},
  journal= {arXiv preprint arXiv:2110.09240},
  year   = {2024}
}

备注

accepted paper at the Responsible Artificial Intelligence Agents Workshop, of the 18th International Conference on Autonomous Agents and MultiAgent Systems (AAMAS 2019)