Perturbed-History Exploration in Stochastic Multi-Armed Bandits
Machine Learning
2019-11-06 v2 Machine Learning
Abstract
We propose an online algorithm for cumulative regret minimization in a stochastic multi-armed bandit. The algorithm adds i.i.d. pseudo-rewards to its history in round and then pulls the arm with the highest average reward in its perturbed history. Therefore, we call it perturbed-history exploration (PHE). The pseudo-rewards are carefully designed to offset potentially underestimated mean rewards of arms with a high probability. We derive near-optimal gap-dependent and gap-free bounds on the -round regret of PHE. The key step in our analysis is a novel argument that shows that randomized Bernoulli rewards lead to optimism. Finally, we empirically evaluate PHE and show that it is competitive with state-of-the-art baselines.
Keywords
Cite
@article{arxiv.1902.10089,
title = {Perturbed-History Exploration in Stochastic Multi-Armed Bandits},
author = {Branislav Kveton and Csaba Szepesvari and Mohammad Ghavamzadeh and Craig Boutilier},
journal= {arXiv preprint arXiv:1902.10089},
year = {2019}
}