中文

多目标强化学习:用于Pluralistic Alignment的工具

机器学习 2024-10-16 v1 人工智能

摘要

强化学习 (RL) 是人工智能系统创建的有价值工具。然而,如果存在多个相互冲突的价值或利益相关者,需要考虑的话,基于标量奖励的足够对齐RL可能存在问题。过去十年来,基于向量奖励的多目标强化学习 (MORL) 作为标准标量 RL 的一种替代方案已逐渐出现。本文提供了MORL在创建Pluralistically对齐的人工智能中所发挥作用的概述。

关键词

引用

@article{arxiv.2410.11221,
  title  = {Multi-objective Reinforcement Learning: A Tool for Pluralistic Alignment},
  author = {Peter Vamplew and Conor F Hayes and Cameron Foale and Richard Dazeley and Hadassah Harland},
  journal= {arXiv preprint arXiv:2410.11221},
  year   = {2024}
}

备注

Accepted for the Pluralistic Alignment workshop at NeurIPS 2024. https://pluralistic-alignment.github.io/