English

Lower Bounds for Multi-armed Bandit with Non-equivalent Multiple Plays

Machine Learning 2015-07-20 v1

Abstract

We study the stochastic multi-armed bandit problem with non-equivalent multiple plays where, at each step, an agent chooses not only a set of arms, but also their order, which influences reward distribution. In several problem formulations with different assumptions, we provide lower bounds for regret with standard asymptotics O(logt)O(\log{t}) but novel coefficients and provide optimal algorithms, thus proving that these bounds cannot be improved.

Keywords

Cite

@article{arxiv.1507.04910,
  title  = {Lower Bounds for Multi-armed Bandit with Non-equivalent Multiple Plays},
  author = {Aleksandr Vorobev and Gleb Gusev},
  journal= {arXiv preprint arXiv:1507.04910},
  year   = {2015}
}
R2 v1 2026-06-22T10:13:48.111Z