中文

Alchemy:一个用于元强化学习智能体的基准与分析工具包

机器学习 2021-10-22 v3 人工智能

摘要

将元学习作为提高强化学习灵活性与样本效率的方法,近来兴趣快速增长。然而该研究领域的一个问题是缺乏充分的基准任务。总体而言,以往基准背后的结构要么过于简单而缺乏内在趣味,要么定义不清而无法支持原则性分析。在本工作中,我们引入一个新的元强化学习(meta-RL)研究基准,强调透明性、深入分析的潜力以及结构丰富性。Alchemy 是一款基于 Unity 实现的 3D 视频游戏,其中包含一集一集程序化重采样的潜在因果结构,支持基于抽象领域知识的结构学习、在线推断、假设检验与动作序列规划。我们在一对强大的 RL 智能体上评估 Alchemy,并对其中之一的智能体进行深入分析。结果清晰地揭示了元学习明确而具体的失败,验证了 Alchemy 作为元强化学习挑战性基准的价值。与本报告同时,我们将 Alchemy 作为公共资源发布,并附带一套分析工具与智能体轨迹样本。

关键词

引用

@article{arxiv.2102.02926,
  title  = {Alchemy: A benchmark and analysis toolkit for meta-reinforcement learning agents},
  author = {Jane X. Wang and Michael King and Nicolas Porcel and Zeb Kurth-Nelson and Tina Zhu and Charlie Deck and Peter Choy and Mary Cassin and Malcolm Reynolds and Francis Song and Gavin Buttimore and David P. Reichert and Neil Rabinowitz and Loic Matthey and Demis Hassabis and Alexander Lerchner and Matthew Botvinick},
  journal= {arXiv preprint arXiv:2102.02926},
  year   = {2021}
}

备注

Published in Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 2021