Unified Framework of Distributional Regret in Multi-Armed Bandits and Reinforcement Learning
Abstract
We study the distribution of regret in stochastic multi-armed bandits and episodic reinforcement learning through a unified framework. We formalize a distributional regret bound as a probabilistic guarantee that holds uniformly over all confidence levels , thereby characterizing the regret distribution across the full range of . We present a simple UCBVI-style algorithm with exploration bonus , where denotes the visit count and are user-specified parameters. For arbitrary parameter sequences, we derive general gap-independent and gap-dependent distributional regret bounds, yielding a principled characterization of how the parameters control the trade-off between expected performance, tail risk, and instance-dependent behavior. In particular, our bounds achieve optimal trade-offs between expected and distributional regret in both minimax and instance-dependent regimes. As a special case, for multi-armed bandits with arms and horizon , we obtain a distributional regret bound of order , confirming the conjecture of Lattimore & Szepesv\'ari (2020, Section 17.1) for the first time.
Keywords
Cite
@article{arxiv.2605.05102,
title = {Unified Framework of Distributional Regret in Multi-Armed Bandits and Reinforcement Learning},
author = {Harin Lee and Min-hwan Oh},
journal= {arXiv preprint arXiv:2605.05102},
year = {2026}
}
Comments
Accepted at the Conference of Learning Theory (COLT) 2026