English

Unified Framework of Distributional Regret in Multi-Armed Bandits and Reinforcement Learning

Machine Learning 2026-05-08 v2 Machine Learning

Abstract

We study the distribution of regret in stochastic multi-armed bandits and episodic reinforcement learning through a unified framework. We formalize a distributional regret bound as a probabilistic guarantee that holds uniformly over all confidence levels δ(0,1]\delta \in (0,1], thereby characterizing the regret distribution across the full range of δ\delta. We present a simple UCBVI-style algorithm with exploration bonus min{c1,k/N,c2,k/N}\min\{c_{1,k}/N, c_{2,k}/\sqrt{N}\}, where NN denotes the visit count and (c1,k,c2,k)(c_{1,k},c_{2,k}) are user-specified parameters. For arbitrary parameter sequences, we derive general gap-independent and gap-dependent distributional regret bounds, yielding a principled characterization of how the parameters control the trade-off between expected performance, tail risk, and instance-dependent behavior. In particular, our bounds achieve optimal trade-offs between expected and distributional regret in both minimax and instance-dependent regimes. As a special case, for multi-armed bandits with AA arms and horizon TT, we obtain a distributional regret bound of order O(ATlog(1/δ))\mathcal{O}(\sqrt{AT}\log(1/\delta)), confirming the conjecture of Lattimore & Szepesv\'ari (2020, Section 17.1) for the first time.

Keywords

Cite

@article{arxiv.2605.05102,
  title  = {Unified Framework of Distributional Regret in Multi-Armed Bandits and Reinforcement Learning},
  author = {Harin Lee and Min-hwan Oh},
  journal= {arXiv preprint arXiv:2605.05102},
  year   = {2026}
}

Comments

Accepted at the Conference of Learning Theory (COLT) 2026