English

Asymptotically optimal regret in communicating Markov decision processes

Machine Learning 2025-05-26 v1 Machine Learning

Abstract

In this paper, we present a learning algorithm that achieves asymptotically optimal regret for Markov decision processes in average reward under a communicating assumption. That is, given a communicating Markov decision process MM, our algorithm has regret K(M)log(T)+o(log(T))K(M) \log(T) + \mathrm{o}(\log(T)) where TT is the number of learning steps and K(M)K(M) is the best possible constant. This algorithm works by explicitly tracking the constant K(M)K(M) to learn optimally, then balances the trade-off between exploration (playing sub-optimally to gain information), co-exploration (playing optimally to gain information) and exploitation (playing optimally to score maximally). We further show that the function K(M)K(M) is discontinuous, which is a consequence challenge for our approach. To that end, we describe a regularization mechanism to estimate K(M)K(M) with arbitrary precision from empirical data.

Keywords

Cite

@article{arxiv.2505.18064,
  title  = {Asymptotically optimal regret in communicating Markov decision processes},
  author = {Victor Boone},
  journal= {arXiv preprint arXiv:2505.18064},
  year   = {2025}
}
R2 v1 2026-07-01T02:34:14.308Z