English

Improved Analysis of UCRL2 with Empirical Bernstein Inequality

Machine Learning 2020-07-13 v1 Machine Learning

Abstract

We consider the problem of exploration-exploitation in communicating Markov Decision Processes. We provide an analysis of UCRL2 with Empirical Bernstein inequalities (UCRL2B). For any MDP with SS states, AA actions, ΓS\Gamma \leq S next states and diameter DD, the regret of UCRL2B is bounded as O~(DΓSAT)\widetilde{O}(\sqrt{D\Gamma S A T}).

Keywords

Cite

@article{arxiv.2007.05456,
  title  = {Improved Analysis of UCRL2 with Empirical Bernstein Inequality},
  author = {Ronan Fruit and Matteo Pirotta and Alessandro Lazaric},
  journal= {arXiv preprint arXiv:2007.05456},
  year   = {2020}
}

Comments

Document in support of the tutorial at ALT 2019