中文

标量奖励并不足够:对 Silver、Singh、Precup 与 Sutton(2021)的回应

人工智能 2022-01-03 v1

摘要

Silver、Singh、Precup 与 Sutton 近期的论文《奖励即足够》主张,奖励最大化这一概念足以支撑所有自然与人工智能。我们质疑 Silver 等人关于此类奖励可为标量值的基本假设。本文中,我们解释为何标量奖励不足以解释生物与计算智能的某些方面,并主张支持显式多目标奖励最大化模型。此外,我们主张即便标量奖励函数能在特定情形下触发智能行为,出于不安全或不道德行为不可接受的风险,该方法仍不应用于开发通用人工智能。

关键词

引用

@article{arxiv.2112.15422,
  title  = {Scalar reward is not enough: A response to Silver, Singh, Precup and Sutton (2021)},
  author = {Peter Vamplew and Benjamin J. Smith and Johan Kallstrom and Gabriel Ramos and Roxana Radulescu and Diederik M. Roijers and Conor F. Hayes and Fredrik Heintz and Patrick Mannion and Pieter J. K. Libin and Richard Dazeley and Cameron Foale},
  journal= {arXiv preprint arXiv:2112.15422},
  year   = {2022}
}