在街机学习环境中对基于奖励 bonuses 的探索方法的基准评测
机器学习
2021-09-28 v3 机器学习
摘要
本文在街机学习环境(ALE)中对近年开发的探索算法进行了实证评估。我们研究了在强化学习中激励探索的不同奖励bonuses的使用。为此,我们固定所采用的学习算法,仅关注不同探索bonuses对智能体性能的影响。我们使用Rainbow——基于价值的智能体的最先进算法——并聚焦于过去几年提出的一些bonuses。我们考虑这些算法在广受探索领域关注的热门游戏Montezuma's Revenge中的性能影响,在Bellemare等人(2016)确定的七款被认定为探索具有挑战性的游戏集合中,以及在探索不成问题的较简单游戏中的影响。我们发现,在我们的设定下,近年开发的bonuses并未在Montezuma's Revenge或困难探索游戏上提供显著改进的性能。我们还发现,现有的基于bonus的方法可能对探索不成问题的游戏性能产生负面影响,甚至可能比-贪婪探索表现更差。
引用
@article{arxiv.1908.02388,
title = {Benchmarking Bonus-Based Exploration Methods on the Arcade Learning Environment},
author = {Adrien Ali Taïga and William Fedus and Marlos C. Machado and Aaron Courville and Marc G. Bellemare},
journal= {arXiv preprint arXiv:1908.02388},
year = {2021}
}
备注
Accepted at the second Exploration in Reinforcement Learning Workshop at the 36th International Conference on Machine Learning, Long Beach, California. The full version arxiv.org/abs/2109.11052 was published as a conference paper at ICLR 2020