English

Bounded regret in stochastic multi-armed bandits

Statistics Theory 2013-02-13 v2 Machine Learning Machine Learning Statistics Theory

Abstract

We study the stochastic multi-armed bandit problem when one knows the value μ()\mu^{(\star)} of an optimal arm, as a well as a positive lower bound on the smallest positive gap Δ\Delta. We propose a new randomized policy that attains a regret {\em uniformly bounded over time} in this setting. We also prove several lower bounds, which show in particular that bounded regret is not possible if one only knows Δ\Delta, and bounded regret of order 1/Δ1/\Delta is not possible if one only knows μ()\mu^{(\star)}

Cite

@article{arxiv.1302.1611,
  title  = {Bounded regret in stochastic multi-armed bandits},
  author = {Sébastien Bubeck and Vianney Perchet and Philippe Rigollet},
  journal= {arXiv preprint arXiv:1302.1611},
  year   = {2013}
}
R2 v1 2026-06-21T23:22:17.544Z