Lower Bounds for Multi-armed Bandit with Non-equivalent Multiple Plays
Machine Learning
2015-07-20 v1
Abstract
We study the stochastic multi-armed bandit problem with non-equivalent multiple plays where, at each step, an agent chooses not only a set of arms, but also their order, which influences reward distribution. In several problem formulations with different assumptions, we provide lower bounds for regret with standard asymptotics but novel coefficients and provide optimal algorithms, thus proving that these bounds cannot be improved.
Cite
@article{arxiv.1507.04910,
title = {Lower Bounds for Multi-armed Bandit with Non-equivalent Multiple Plays},
author = {Aleksandr Vorobev and Gleb Gusev},
journal= {arXiv preprint arXiv:1507.04910},
year = {2015}
}