English

Observe Before Play: Multi-armed Bandit with Pre-observations

Machine Learning 2019-11-22 v1 Machine Learning

Abstract

We consider the stochastic multi-armed bandit (MAB) problem in a setting where a player can pay to pre-observe arm rewards before playing an arm in each round. Apart from the usual trade-off between exploring new arms to find the best one and exploiting the arm believed to offer the highest reward, we encounter an additional dilemma: pre-observing more arms gives a higher chance to play the best one, but incurs a larger cost. For the single-player setting, we design an Observe-Before-Play Upper Confidence Bound (OBP-UCB) algorithm for KK arms with Bernoulli rewards, and prove a TT-round regret upper bound O(K2logT)O(K^2\log T). In the multi-player setting, collisions will occur when players select the same arm to play in the same round. We design a centralized algorithm, C-MP-OBP, and prove its TT-round regret relative to an offline greedy strategy is upper bounded in O(K4M2logT)O(\frac{K^4}{M^2}\log T) for KK arms and MM players. We also propose distributed versions of the C-MP-OBP policy, called D-MP-OBP and D-MP-Adapt-OBP, achieving logarithmic regret with respect to collision-free target policies. Experiments on synthetic data and wireless channel traces show that C-MP-OBP and D-MP-OBP outperform random heuristics and offline optimal policies that do not allow pre-observations.

Keywords

Cite

@article{arxiv.1911.09458,
  title  = {Observe Before Play: Multi-armed Bandit with Pre-observations},
  author = {Jinhang Zuo and Xiaoxi Zhang and Carlee Joe-Wong},
  journal= {arXiv preprint arXiv:1911.09458},
  year   = {2019}
}
R2 v1 2026-06-23T12:23:20.845Z