基于经验 Bernstein 不等式的 UCRL2 改进分析
机器学习
2020-07-13 v1 机器学习
摘要
我们考虑通信马尔可夫决策过程中的探索-利用问题。我们提供了基于经验 Bernstein 不等式的 UCRL2 分析(UCRL2B)。对于任意具有 个状态、 个动作、 个后继状态以及直径 的 MDP,UCRL2B 的遗憾界为 。
引用
@article{arxiv.2007.05456,
title = {Improved Analysis of UCRL2 with Empirical Bernstein Inequality},
author = {Ronan Fruit and Matteo Pirotta and Alessandro Lazaric},
journal= {arXiv preprint arXiv:2007.05456},
year = {2020}
}
备注
Document in support of the tutorial at ALT 2019