Improved Analysis of UCRL2 with Empirical Bernstein Inequality
Machine Learning
2020-07-13 v1 Machine Learning
Abstract
We consider the problem of exploration-exploitation in communicating Markov Decision Processes. We provide an analysis of UCRL2 with Empirical Bernstein inequalities (UCRL2B). For any MDP with states, actions, next states and diameter , the regret of UCRL2B is bounded as .
Keywords
Cite
@article{arxiv.2007.05456,
title = {Improved Analysis of UCRL2 with Empirical Bernstein Inequality},
author = {Ronan Fruit and Matteo Pirotta and Alessandro Lazaric},
journal= {arXiv preprint arXiv:2007.05456},
year = {2020}
}
Comments
Document in support of the tutorial at ALT 2019