English

Planning in entropy-regularized Markov decision processes and games

Machine Learning 2026-04-22 v1

Abstract

We propose SmoothCruiser, a new planning algorithm for estimating the value function in entropy-regularized Markov decision processes and two-player games, given a generative model of the environment. SmoothCruiser makes use of the smoothness of the Bellman operator promoted by the regularization to achieve problem-independent sample complexity of order O~(1/epsilon^4) for a desired accuracy epsilon, whereas for non-regularized settings there are no known algorithms with guaranteed polynomial sample complexity in the worst case.

Keywords

Cite

@article{arxiv.2604.19695,
  title  = {Planning in entropy-regularized Markov decision processes and games},
  author = {Jean-Bastien Grill and Omar Darwiche Domingues and Pierre Ménard and Rémi Munos and Michal Valko},
  journal= {arXiv preprint arXiv:2604.19695},
  year   = {2026}
}

Comments

NeurIPS 2019

R2 v1 2026-07-01T12:28:48.443Z