English

Analysis of Thompson Sampling for Combinatorial Multi-armed Bandit with Probabilistically Triggered Arms

Machine Learning 2019-02-20 v2 Machine Learning

Abstract

We analyze the regret of combinatorial Thompson sampling (CTS) for the combinatorial multi-armed bandit with probabilistically triggered arms under the semi-bandit feedback setting. We assume that the learner has access to an exact optimization oracle but does not know the expected base arm outcomes beforehand. When the expected reward function is Lipschitz continuous in the expected base arm outcomes, we derive O(i=1mlogT/(piΔi))O(\sum_{i =1}^m \log T / (p_i \Delta_i)) regret bound for CTS, where mm denotes the number of base arms, pip_i denotes the minimum non-zero triggering probability of base arm ii and Δi\Delta_i denotes the minimum suboptimality gap of base arm ii. We also compare CTS with combinatorial upper confidence bound (CUCB) via numerical experiments on a cascading bandit problem.

Keywords

Cite

@article{arxiv.1809.02707,
  title  = {Analysis of Thompson Sampling for Combinatorial Multi-armed Bandit with Probabilistically Triggered Arms},
  author = {Alihan Hüyük and Cem Tekin},
  journal= {arXiv preprint arXiv:1809.02707},
  year   = {2019}
}

Comments

To appear in the Proceedings of the 22nd International Conference on Artificial Intelligence and Statistics (AISTATS) 2019

R2 v1 2026-06-23T03:58:37.459Z