Nearly Minimax Optimal Reinforcement Learning for Discounted MDPs
Machine Learning
2022-01-04 v3 Optimization and Control
Machine Learning
Abstract
We study the reinforcement learning problem for discounted Markov Decision Processes (MDPs) under the tabular setting. We propose a model-based algorithm named UCBVI-, which is based on the \emph{optimism in the face of uncertainty principle} and the Bernstein-type bonus. We show that UCBVI- achieves an regret, where is the number of states, is the number of actions, is the discount factor and is the number of steps. In addition, we construct a class of hard MDPs and show that for any algorithm, the expected regret is at least . Our upper bound matches the minimax lower bound up to logarithmic factors, which suggests that UCBVI- is nearly minimax optimal for discounted MDPs.
Cite
@article{arxiv.2010.00587,
title = {Nearly Minimax Optimal Reinforcement Learning for Discounted MDPs},
author = {Jiafan He and Dongruo Zhou and Quanquan Gu},
journal= {arXiv preprint arXiv:2010.00587},
year = {2022}
}
Comments
33 pages, 1 figure, 1 table. In NeurIPS 2021