Adaptation to the Range in $K$-Armed Bandits
Statistics Theory
2022-06-16 v3 Machine Learning
Statistics Theory
Abstract
We consider stochastic bandit problems with arms, each associated with a bounded distribution supported on the range . We do not assume that the range is known and show that there is a cost for learning this range. Indeed, a new trade-off between distribution-dependent and distribution-free regret bounds arises, which prevents from simultaneously achieving the typical and bounds. For instance, a }distribution-free regret bound may only be achieved if the distribution-dependent regret bounds are at least of order . We exhibit a strategy achieving the rates for regret indicated by the new trade-off.
Cite
@article{arxiv.2006.03378,
title = {Adaptation to the Range in $K$-Armed Bandits},
author = {Hédi Hadiji and Gilles Stoltz},
journal= {arXiv preprint arXiv:2006.03378},
year = {2022}
}