English

Provably Adaptive Average Reward Reinforcement Learning for Metric Spaces

Machine Learning 2026-02-03 v2 Artificial Intelligence

Abstract

We study infinite-horizon average-reward reinforcement learning (RL) for Lipschitz MDPs, a broad class that subsumes several important classes such as linear and RKHS MDPs, function approximation frameworks, and develop an adaptive algorithm ZoRL\text{ZoRL} with regret bounded as O(T1deff.1)\mathcal{O}\big(T^{1 - d_{\text{eff.}}^{-1}}\big), where deff.=2dS+dz+3d_{\text{eff.}}= 2d_\mathcal{S} + d_z + 3, dSd_\mathcal{S} is the dimension of the state space and dzd_z is the zooming dimension. In contrast, algorithms with fixed discretization yield deff.=2(dS+dA)+2d_{\text{eff.}} = 2(d_\mathcal{S} + d_\mathcal{A}) + 2, dAd_\mathcal{A} being the dimension of action space. ZoRL\text{ZoRL} achieves this by discretizing the state-action space adaptively and zooming into ''promising regions'' of the state-action space. dzd_z, a problem-dependent quantity bounded by the state-action space's dimension, allows us to conclude that if an MDP is benign, then the regret of ZoRL\text{ZoRL} will be small. The zooming dimension and ZoRL\text{ZoRL} are truly adaptive, i.e., the current work shows how to capture adaptivity gains for infinite-horizon average-reward RL. ZoRL\text{ZoRL} outperforms other state-of-the-art algorithms in experiments, thereby demonstrating the gains arising due to adaptivity.

Keywords

Cite

@article{arxiv.2410.19919,
  title  = {Provably Adaptive Average Reward Reinforcement Learning for Metric Spaces},
  author = {Avik Kar and Rahul Singh},
  journal= {arXiv preprint arXiv:2410.19919},
  year   = {2026}
}

Comments

Accepted in the 41st Conference on Uncertainty in Artificial Intelligence

R2 v1 2026-06-28T19:36:08.129Z