中文

PAC 模型无关强化学习中的定向探索

机器学习 2018-09-03 v1 机器学习

摘要

我们研究一种模型无关强化学习(RL)的探索方法,它推广了基于计数的探索奖励方法,并考虑动作的长期探索价值而非单步前瞻。我们提出一种模型无关 RL 方法,改进了 Delayed Q-learning 并利用具有可证明效率的长期探索奖励。我们证明所提方法在多项式时间内(PAC-MDP)找到近最优策略,并提供实验证据表明所提算法是一种高效的探索方法。

关键词

引用

@article{arxiv.1808.10552,
  title  = {Directed Exploration in PAC Model-Free Reinforcement Learning},
  author = {Min-hwan Oh and Garud Iyengar},
  journal= {arXiv preprint arXiv:1808.10552},
  year   = {2018}
}