English

Explore First, Exploit Next: The True Shape of Regret in Bandit Problems

Statistics Theory 2018-10-16 v3 Machine Learning Statistics Theory

Abstract

We revisit lower bounds on the regret in the case of multi-armed bandit problems. We obtain non-asymptotic, distribution-dependent bounds and provide straightforward proofs based only on well-known properties of Kullback-Leibler divergences. These bounds show in particular that in an initial phase the regret grows almost linearly, and that the well-known logarithmic growth of the regret only holds in a final phase. The proof techniques come to the essence of the information-theoretic arguments used and they are deprived of all unnecessary complications.

Cite

@article{arxiv.1602.07182,
  title  = {Explore First, Exploit Next: The True Shape of Regret in Bandit Problems},
  author = {Aurélien Garivier and Pierre Ménard and Gilles Stoltz},
  journal= {arXiv preprint arXiv:1602.07182},
  year   = {2018}
}
R2 v1 2026-06-22T12:56:01.522Z