English

Decentralized Heterogeneous Multi-Player Multi-Armed Bandits with Non-Zero Rewards on Collisions

Machine Learning 2021-12-30 v4 Computer Science and Game Theory Machine Learning

Abstract

We consider a fully decentralized multi-player stochastic multi-armed bandit setting where the players cannot communicate with each other and can observe only their own actions and rewards. The environment may appear differently to different players, i.e.\textit{i.e.}, the reward distributions for a given arm are heterogeneous across players. In the case of a collision (when more than one player plays the same arm), we allow for the colliding players to receive non-zero rewards. The time-horizon TT for which the arms are played is \emph{not} known to the players. Within this setup, where the number of players is allowed to be greater than the number of arms, we present a policy that achieves near order-optimal expected regret of order O(log1+δT)O(\log^{1 + \delta} T) for some 0<δ<10 < \delta < 1 over a time-horizon of duration TT. This paper is accepted at IEEE Transactions on Information Theory.

Keywords

Cite

@article{arxiv.1910.09089,
  title  = {Decentralized Heterogeneous Multi-Player Multi-Armed Bandits with Non-Zero Rewards on Collisions},
  author = {Akshayaa Magesh and Venugopal V. Veeravalli},
  journal= {arXiv preprint arXiv:1910.09089},
  year   = {2021}
}