English

Non-Stochastic Multi-Player Multi-Armed Bandits: Optimal Rate With Collision Information, Sublinear Without

Machine Learning 2019-05-03 v2 Multiagent Systems Machine Learning

Abstract

We consider the non-stochastic version of the (cooperative) multi-player multi-armed bandit problem. The model assumes no communication at all between the players, and furthermore when two (or more) players select the same action this results in a maximal loss. We prove the first T\sqrt{T}-type regret guarantee for this problem, under the feedback model where collisions are announced to the colliding players. Such a bound was not known even for the simpler stochastic version. We also prove the first sublinear guarantee for the feedback model where collision information is not available, namely T112mT^{1-\frac{1}{2m}} where mm is the number of players.

Keywords

Cite

@article{arxiv.1904.12233,
  title  = {Non-Stochastic Multi-Player Multi-Armed Bandits: Optimal Rate With Collision Information, Sublinear Without},
  author = {Sébastien Bubeck and Yuanzhi Li and Yuval Peres and Mark Sellke},
  journal= {arXiv preprint arXiv:1904.12233},
  year   = {2019}
}

Comments

27 pages, v2 adds a pseudorandom generator construction to remove the shared randomness assumption in the $\sqrt{T}$-regret result (Section 3.9)

R2 v1 2026-06-23T08:51:21.064Z