English

On Distributed Cooperative Decision-Making in Multiarmed Bandits

Systems and Control 2019-09-17 v3 Multiagent Systems Optimization and Control Machine Learning

Abstract

We study the explore-exploit tradeoff in distributed cooperative decision-making using the context of the multiarmed bandit (MAB) problem. For the distributed cooperative MAB problem, we design the cooperative UCB algorithm that comprises two interleaved distributed processes: (i) running consensus algorithms for estimation of rewards, and (ii) upper-confidence-bound-based heuristics for selection of arms. We rigorously analyze the performance of the cooperative UCB algorithm and characterize the influence of communication graph structure on the decision-making performance of the group.

Keywords

Cite

@article{arxiv.1512.06888,
  title  = {On Distributed Cooperative Decision-Making in Multiarmed Bandits},
  author = {Peter Landgren and Vaibhav Srivastava and Naomi Ehrich Leonard},
  journal= {arXiv preprint arXiv:1512.06888},
  year   = {2019}
}

Comments

This revision provides a correction to the original paper, which appeared in the Proceedings of the 2016 European Control Conference (ECC). The second statement of Proposition 1, Theorem 1 and their proofs are new. The new Theorem 1 is used to prove the regret bounds in Theorem 2

R2 v1 2026-06-22T12:15:28.672Z