On Distributed Cooperative Decision-Making in Multiarmed Bandits
Abstract
We study the explore-exploit tradeoff in distributed cooperative decision-making using the context of the multiarmed bandit (MAB) problem. For the distributed cooperative MAB problem, we design the cooperative UCB algorithm that comprises two interleaved distributed processes: (i) running consensus algorithms for estimation of rewards, and (ii) upper-confidence-bound-based heuristics for selection of arms. We rigorously analyze the performance of the cooperative UCB algorithm and characterize the influence of communication graph structure on the decision-making performance of the group.
Cite
@article{arxiv.1512.06888,
title = {On Distributed Cooperative Decision-Making in Multiarmed Bandits},
author = {Peter Landgren and Vaibhav Srivastava and Naomi Ehrich Leonard},
journal= {arXiv preprint arXiv:1512.06888},
year = {2019}
}
Comments
This revision provides a correction to the original paper, which appeared in the Proceedings of the 2016 European Control Conference (ECC). The second statement of Proposition 1, Theorem 1 and their proofs are new. The new Theorem 1 is used to prove the regret bounds in Theorem 2