中文

回报分布在强化学习探索中的潜力

机器学习 2018-07-04 v2 人工智能 机器学习

摘要

本文研究在确定性强化学习(RL)环境中,回报分布用于探索的潜力。我们研究 Gaussian、Categorical 与 Gaussian mixture 分布的网络损失与传播机制。结合利用该回报分布的探索策略,我们求解了例如长度为 100 的随机 Chain 任务,这是使用神经网络学习时此前未见报道的。

关键词

引用

@article{arxiv.1806.04242,
  title  = {The Potential of the Return Distribution for Exploration in RL},
  author = {Thomas M. Moerland and Joost Broekens and Catholijn M. Jonker},
  journal= {arXiv preprint arXiv:1806.04242},
  year   = {2018}
}

备注

Published at the Exploration in Reinforcement Learning Workshop at the 35th International Conference on Machine Learning, Stockholm, Sweden