English

KL-UCB-switch: optimal regret bounds for stochastic bandits from both a distribution-dependent and a distribution-free viewpoints

Machine Learning 2022-07-04 v3 Machine Learning Statistics Theory Statistics Theory

Abstract

We consider KK-armed stochastic bandits and consider cumulative regret bounds up to time TT. We are interested in strategies achieving simultaneously a distribution-free regret bound of optimal order KT\sqrt{KT} and a distribution-dependent regret that is asymptotically optimal, that is, matching the κlnT\kappa\ln T lower bound by Lai and Robbins (1985) and Burnetas and Katehakis (1996), where κ\kappa is the optimal problem-dependent constant. This constant κ\kappa depends on the model D\mathcal{D} considered (the family of possible distributions over the arms). M\'enard and Garivier (2017) provided strategies achieving such a bi-optimality in the parametric case of models given by one-dimensional exponential families, while Lattimore (2016, 2018) did so for the family of (sub)Gaussian distributions with variance less than 11. We extend this result to the non-parametric case of all distributions over [0,1][0,1]. We do so by combining the MOSS strategy by Audibert and Bubeck (2009), which enjoys a distribution-free regret bound of optimal order KT\sqrt{KT}, and the KL-UCB strategy by Capp\'e et al. (2013), for which we provide in passing the first analysis of an optimal distribution-dependent κlnT\kappa\ln T regret bound in the model of all distributions over [0,1][0,1]. We were able to obtain this non-parametric bi-optimality result while working hard to streamline the proofs (of previously known regret bounds and thus of the new analyses carried out); a second merit of the present contribution is therefore to provide a review of proofs of classical regret bounds for index-based strategies for KK-armed stochastic bandits.

Keywords

Cite

@article{arxiv.1805.05071,
  title  = {KL-UCB-switch: optimal regret bounds for stochastic bandits from both a distribution-dependent and a distribution-free viewpoints},
  author = {Aurélien Garivier and Hédi Hadiji and Pierre Menard and Gilles Stoltz},
  journal= {arXiv preprint arXiv:1805.05071},
  year   = {2022}
}