Asymptotically optimal regret in communicating Markov decision processes
Abstract
In this paper, we present a learning algorithm that achieves asymptotically optimal regret for Markov decision processes in average reward under a communicating assumption. That is, given a communicating Markov decision process , our algorithm has regret where is the number of learning steps and is the best possible constant. This algorithm works by explicitly tracking the constant to learn optimally, then balances the trade-off between exploration (playing sub-optimally to gain information), co-exploration (playing optimally to gain information) and exploitation (playing optimally to score maximally). We further show that the function is discontinuous, which is a consequence challenge for our approach. To that end, we describe a regularization mechanism to estimate with arbitrary precision from empirical data.
Keywords
Cite
@article{arxiv.2505.18064,
title = {Asymptotically optimal regret in communicating Markov decision processes},
author = {Victor Boone},
journal= {arXiv preprint arXiv:2505.18064},
year = {2025}
}