Bounded regret in stochastic multi-armed bandits
Statistics Theory
2013-02-13 v2 Machine Learning
Machine Learning
Statistics Theory
Abstract
We study the stochastic multi-armed bandit problem when one knows the value of an optimal arm, as a well as a positive lower bound on the smallest positive gap . We propose a new randomized policy that attains a regret {\em uniformly bounded over time} in this setting. We also prove several lower bounds, which show in particular that bounded regret is not possible if one only knows , and bounded regret of order is not possible if one only knows
Cite
@article{arxiv.1302.1611,
title = {Bounded regret in stochastic multi-armed bandits},
author = {Sébastien Bubeck and Vianney Perchet and Philippe Rigollet},
journal= {arXiv preprint arXiv:1302.1611},
year = {2013}
}