Coordination without communication: optimal regret in two players multi-armed bandits
Computer Science and Game Theory
2020-07-13 v2 Machine Learning
Multiagent Systems
Machine Learning
Abstract
We consider two agents playing simultaneously the same stochastic three-armed bandit problem. The two agents are cooperating but they cannot communicate. We propose a strategy with no collisions at all between the players (with very high probability), and with near-optimal regret . We also argue that the extra logarithmic term should be necessary by proving a lower bound for a full information variant of the problem.
Keywords
Cite
@article{arxiv.2002.07596,
title = {Coordination without communication: optimal regret in two players multi-armed bandits},
author = {Sébastien Bubeck and Thomas Budzinski},
journal= {arXiv preprint arXiv:2002.07596},
year = {2020}
}
Comments
28 pages, 5 figures. V2: minor revision