English

Coordination without communication: optimal regret in two players multi-armed bandits

Computer Science and Game Theory 2020-07-13 v2 Machine Learning Multiagent Systems Machine Learning

Abstract

We consider two agents playing simultaneously the same stochastic three-armed bandit problem. The two agents are cooperating but they cannot communicate. We propose a strategy with no collisions at all between the players (with very high probability), and with near-optimal regret O(Tlog(T))O(\sqrt{T \log(T)}). We also argue that the extra logarithmic term log(T)\sqrt{\log(T)} should be necessary by proving a lower bound for a full information variant of the problem.

Keywords

Cite

@article{arxiv.2002.07596,
  title  = {Coordination without communication: optimal regret in two players multi-armed bandits},
  author = {Sébastien Bubeck and Thomas Budzinski},
  journal= {arXiv preprint arXiv:2002.07596},
  year   = {2020}
}

Comments

28 pages, 5 figures. V2: minor revision

R2 v1 2026-06-23T13:45:23.556Z