English

Tail Distribution of Regret in Optimistic Reinforcement Learning

Machine Learning 2026-03-18 v3 Optimization and Control

Abstract

We derive instance-dependent tail bounds for the regret of optimism-based reinforcement learning in finite-horizon tabular Markov decision processes with unknown transition dynamics. We first study a UCBVI-type (model-based) algorithm and characterize the tail distribution of the cumulative regret RKR_K over KK episodes via explicit bounds on P(RKx)P(R_K \ge x), going beyond analyses limited to E[RK]E[R_K] or a single high-probability quantile. We analyze two natural exploration-bonus schedules for UCBVI: (i) a KK-dependent scheme that explicitly incorporates the total number of episodes KK, and (ii) a KK-independent (anytime) scheme that depends only on the current episode index. We then complement the model-based results with an analysis of optimistic Q-learning (model-free) under a KK-dependent bonus schedule. Across both the model-based and model-free settings, we obtain upper bounds on P(RKx)P(R_K \ge x) with a distinctive two-regime structure: a sub-Gaussian tail starting from an instance-dependent scale up to a transition threshold, followed by a sub-Weibull tail beyond that point. We further derive corresponding instance-dependent bounds on the expected regret E[RK]E[R_K]. The proposed algorithms depend on a tuning parameter α\alpha, which balances the expected regret and the range over which the regret exhibits sub-Gaussian decay.

Keywords

Cite

@article{arxiv.2511.18247,
  title  = {Tail Distribution of Regret in Optimistic Reinforcement Learning},
  author = {Sajad Khodadadian and Mehrdad Moharrami},
  journal= {arXiv preprint arXiv:2511.18247},
  year   = {2026}
}

Comments

27 pages, 0 figures