中文

基于经验 Bernstein 不等式的 UCRL2 改进分析

机器学习 2020-07-13 v1 机器学习

摘要

我们考虑通信马尔可夫决策过程中的探索-利用问题。我们提供了基于经验 Bernstein 不等式的 UCRL2 分析(UCRL2B)。对于任意具有 SS 个状态、AA 个动作、ΓS\Gamma \leq S 个后继状态以及直径 DD 的 MDP,UCRL2B 的遗憾界为 O~(DΓSAT)\widetilde{O}(\sqrt{D\Gamma S A T})

关键词

引用

@article{arxiv.2007.05456,
  title  = {Improved Analysis of UCRL2 with Empirical Bernstein Inequality},
  author = {Ronan Fruit and Matteo Pirotta and Alessandro Lazaric},
  journal= {arXiv preprint arXiv:2007.05456},
  year   = {2020}
}

备注

Document in support of the tutorial at ALT 2019