Planning in entropy-regularized Markov decision processes and games
Machine Learning
2026-04-22 v1
Abstract
We propose SmoothCruiser, a new planning algorithm for estimating the value function in entropy-regularized Markov decision processes and two-player games, given a generative model of the environment. SmoothCruiser makes use of the smoothness of the Bellman operator promoted by the regularization to achieve problem-independent sample complexity of order O~(1/epsilon^4) for a desired accuracy epsilon, whereas for non-regularized settings there are no known algorithms with guaranteed polynomial sample complexity in the worst case.
Keywords
Cite
@article{arxiv.2604.19695,
title = {Planning in entropy-regularized Markov decision processes and games},
author = {Jean-Bastien Grill and Omar Darwiche Domingues and Pierre Ménard and Rémi Munos and Michal Valko},
journal= {arXiv preprint arXiv:2604.19695},
year = {2026}
}
Comments
NeurIPS 2019