Provably Adaptive Average Reward Reinforcement Learning for Metric Spaces
Abstract
We study infinite-horizon average-reward reinforcement learning (RL) for Lipschitz MDPs, a broad class that subsumes several important classes such as linear and RKHS MDPs, function approximation frameworks, and develop an adaptive algorithm with regret bounded as , where , is the dimension of the state space and is the zooming dimension. In contrast, algorithms with fixed discretization yield , being the dimension of action space. achieves this by discretizing the state-action space adaptively and zooming into ''promising regions'' of the state-action space. , a problem-dependent quantity bounded by the state-action space's dimension, allows us to conclude that if an MDP is benign, then the regret of will be small. The zooming dimension and are truly adaptive, i.e., the current work shows how to capture adaptivity gains for infinite-horizon average-reward RL. outperforms other state-of-the-art algorithms in experiments, thereby demonstrating the gains arising due to adaptivity.
Keywords
Cite
@article{arxiv.2410.19919,
title = {Provably Adaptive Average Reward Reinforcement Learning for Metric Spaces},
author = {Avik Kar and Rahul Singh},
journal= {arXiv preprint arXiv:2410.19919},
year = {2026}
}
Comments
Accepted in the 41st Conference on Uncertainty in Artificial Intelligence