English

No-Regret Reinforcement Learning in Smooth MDPs

Machine Learning 2024-02-07 v1 Artificial Intelligence

Abstract

Obtaining no-regret guarantees for reinforcement learning (RL) in the case of problems with continuous state and/or action spaces is still one of the major open challenges in the field. Recently, a variety of solutions have been proposed, but besides very specific settings, the general problem remains unsolved. In this paper, we introduce a novel structural assumption on the Markov decision processes (MDPs), namely ν\nu-smoothness, that generalizes most of the settings proposed so far (e.g., linear MDPs and Lipschitz MDPs). To face this challenging scenario, we propose two algorithms for regret minimization in ν\nu-smooth MDPs. Both algorithms build upon the idea of constructing an MDP representation through an orthogonal feature map based on Legendre polynomials. The first algorithm, \textsc{Legendre-Eleanor}, archives the no-regret property under weaker assumptions but is computationally inefficient, whereas the second one, \textsc{Legendre-LSVI}, runs in polynomial time, although for a smaller class of problems. After analyzing their regret properties, we compare our results with state-of-the-art ones from RL theory, showing that our algorithms achieve the best guarantees.

Keywords

Cite

@article{arxiv.2402.03792,
  title  = {No-Regret Reinforcement Learning in Smooth MDPs},
  author = {Davide Maran and Alberto Maria Metelli and Matteo Papini and Marcello Restell},
  journal= {arXiv preprint arXiv:2402.03792},
  year   = {2024}
}
R2 v1 2026-06-28T14:39:48.754Z