English

Adaptation to the Range in $K$-Armed Bandits

Statistics Theory 2022-06-16 v3 Machine Learning Statistics Theory

Abstract

We consider stochastic bandit problems with KK arms, each associated with a bounded distribution supported on the range [m,M][m,M]. We do not assume that the range [m,M][m,M] is known and show that there is a cost for learning this range. Indeed, a new trade-off between distribution-dependent and distribution-free regret bounds arises, which prevents from simultaneously achieving the typical lnT\ln T and T\sqrt{T} bounds. For instance, a T\sqrt{T}}distribution-free regret bound may only be achieved if the distribution-dependent regret bounds are at least of order T\sqrt{T}. We exhibit a strategy achieving the rates for regret indicated by the new trade-off.

Cite

@article{arxiv.2006.03378,
  title  = {Adaptation to the Range in $K$-Armed Bandits},
  author = {Hédi Hadiji and Gilles Stoltz},
  journal= {arXiv preprint arXiv:2006.03378},
  year   = {2022}
}
R2 v1 2026-06-23T16:05:10.587Z