Stochastic Multi-armed Bandits in Constant Space
Data Structures and Algorithms
2018-05-17 v2 Machine Learning
Machine Learning
Abstract
We consider the stochastic bandit problem in the sublinear space setting, where one cannot record the win-loss record for all arms. We give an algorithm using words of space with regret where is the gap between the best arm and arm and is the gap between the best and the second-best arms. If the rewards are bounded away from and , this is within an factor of the optimum regret possible without space constraints.
Cite
@article{arxiv.1712.09007,
title = {Stochastic Multi-armed Bandits in Constant Space},
author = {David Liau and Eric Price and Zhao Song and Ger Yang},
journal= {arXiv preprint arXiv:1712.09007},
year = {2018}
}
Comments
AISTATS 2018