回报分布在强化学习探索中的潜力
机器学习
2018-07-04 v2 人工智能
机器学习
摘要
本文研究在确定性强化学习(RL)环境中,回报分布用于探索的潜力。我们研究 Gaussian、Categorical 与 Gaussian mixture 分布的网络损失与传播机制。结合利用该回报分布的探索策略,我们求解了例如长度为 100 的随机 Chain 任务,这是使用神经网络学习时此前未见报道的。
引用
@article{arxiv.1806.04242,
title = {The Potential of the Return Distribution for Exploration in RL},
author = {Thomas M. Moerland and Joost Broekens and Catholijn M. Jonker},
journal= {arXiv preprint arXiv:1806.04242},
year = {2018}
}
备注
Published at the Exploration in Reinforcement Learning Workshop at the 35th International Conference on Machine Learning, Stockholm, Sweden