English

Optimal Variance-Dependent Regret Bounds for Infinite-Horizon MDPs

Machine Learning 2026-03-26 v1 Information Theory math.IT Optimization and Control Machine Learning

Abstract

Online reinforcement learning in infinite-horizon Markov decision processes (MDPs) remains less theoretically and algorithmically developed than its episodic counterpart, with many algorithms suffering from high ``burn-in'' costs and failing to adapt to benign instance-specific complexity. In this work, we address these shortcomings for two infinite-horizon objectives: the classical average-reward regret and the γ\gamma-regret. We develop a single tractable UCB-style algorithm applicable to both settings, which achieves the first optimal variance-dependent regret guarantees. Our regret bounds in both settings take the form O~(SAVar+lower-order terms)\tilde{O}( \sqrt{SA\,\text{Var}} + \text{lower-order terms}), where S,AS,A are the state and action space sizes, and Var\text{Var} captures cumulative transition variance. This implies minimax-optimal average-reward and γ\gamma-regret bounds in the worst case but also adapts to easier problem instances, for example yielding nearly constant regret in deterministic MDPs. Furthermore, our algorithm enjoys significantly improved lower-order terms for the average-reward setting. With prior knowledge of the optimal bias span hsp\Vert h^\star\Vert_\text{sp}, our algorithm obtains lower-order terms scaling as hspS2A\Vert h^\star\Vert_\text{sp} S^2 A, which we prove is optimal in both hsp\Vert h^\star\Vert_\text{sp} and AA. Without prior knowledge, we prove that no algorithm can have lower-order terms smaller than hsp2SA\Vert h^\star \Vert_\text{sp}^2 S A, and we provide a prior-free algorithm whose lower-order terms scale as hsp2S3A\Vert h^\star\Vert_\text{sp}^2 S^3 A, nearly matching this lower bound. Taken together, these results completely characterize the optimal dependence on hsp\Vert h^\star\Vert_\text{sp} in both leading and lower-order terms, and reveal a fundamental gap in what is achievable with and without prior knowledge.

Keywords

Cite

@article{arxiv.2603.23926,
  title  = {Optimal Variance-Dependent Regret Bounds for Infinite-Horizon MDPs},
  author = {Guy Zamir and Matthew Zurek and Yudong Chen},
  journal= {arXiv preprint arXiv:2603.23926},
  year   = {2026}
}
R2 v1 2026-07-01T11:36:42.126Z