SafeLife 1.0:在复杂环境中探索副作用
人工智能
2021-03-01 v2
摘要
我们提出 SafeLife,一个公开可用的强化学习环境,用于测试强化学习智能体的安全性。它包含复杂、动态、可调、程序化生成的关卡,其中存在许多不安全行为的机会。智能体既根据其最大化显式奖励的能力评分,也根据其在不产生不必要副作用的情况下安全操作的能力评分。我们使用近端策略优化训练智能体以最大化奖励,并在一组基准关卡上对他们评分。所得智能体性能良好但不安全——它们往往在其环境中造成大的副作用——但它们构成了未来安全性研究可据以衡量的基线。
引用
@article{arxiv.1912.01217,
title = {SafeLife 1.0: Exploring Side Effects in Complex Environments},
author = {Carroll L. Wainwright and Peter Eckersley},
journal= {arXiv preprint arXiv:1912.01217},
year = {2021}
}
备注
Updated version was presented at the AAAI SafeAI 2020 Workshop, but now with updated contact info. Previously presented at the 2019 NeurIPS Safety and Robustness in Decision Making Workshop