关于 Atari 2600 游戏中的灾难性干扰
机器学习
2020-06-11 v2 人工智能
机器学习
摘要
无模型深度强化学习样本效率低下。一个被推测但未经证实的假设是,环境内的灾难性干扰抑制了学习。我们通过街机学习环境(ALE)中的大规模实证研究检验了这一假设,并确实发现了支持性证据。我们表明干扰导致性能进入平台期;网络无法在不降低用于到达该平台期的策略的情况下,对平台期之后的片段进行训练。通过合成控制干扰,我们展示了跨架构、学习算法和环境的性能提升。更精细的分析表明,学习游戏的某一片段往往会提高其他地方的预测误差。我们的研究在强化学习中提供了灾难性干扰与样本效率之间清晰的实证联系。
引用
@article{arxiv.2002.12499,
title = {On Catastrophic Interference in Atari 2600 Games},
author = {William Fedus and Dibya Ghosh and John D. Martin and Marc G. Bellemare and Yoshua Bengio and Hugo Larochelle},
journal= {arXiv preprint arXiv:2002.12499},
year = {2020}
}
备注
First two authors contributed equally. Code available to reproduce experiments at https://github.com/google-research/google-research/tree/master/memento