Bounded Regret for Finite-Armed Structured Bandits
Machine Learning
2014-11-12 v1
Abstract
We study a new type of K-armed bandit problem where the expected return of one arm may depend on the returns of other arms. We present a new algorithm for this general class of problems and show that under certain circumstances it is possible to achieve finite expected cumulative regret. We also give problem-dependent lower bounds on the cumulative regret showing that at least in special cases the new algorithm is nearly optimal.
Cite
@article{arxiv.1411.2919,
title = {Bounded Regret for Finite-Armed Structured Bandits},
author = {Tor Lattimore and Remi Munos},
journal= {arXiv preprint arXiv:1411.2919},
year = {2014}
}
Comments
16 pages