中文

具有概率触发臂的组合多臂_bandit中Thompson采样分析

机器学习 2019-02-20 v2 机器学习

摘要

我们在半bandit反馈设定下,分析了具有概率触发臂的组合多臂bandit中组合Thompson采样(CTS)的遗憾值。我们假设学习器可以访问一个精确优化预言机,但事先不知道期望基臂结果。当期望奖励函数在期望基臂结果上Lipschitz连续时,我们推导出CTS的O(i=1mlogT/(piΔi))O(\sum_{i =1}^m \log T / (p_i \Delta_i))遗憾界,其中mm表示基臂的数量,pip_i表示基臂ii的最小非零触发概率,Δi\Delta_i表示基臂ii的最小次优间隙。我们还通过在级联bandit问题上的数值实验将CTS与组合上置信界(CUCB)进行比较。

关键词

引用

@article{arxiv.1809.02707,
  title  = {Analysis of Thompson Sampling for Combinatorial Multi-armed Bandit with Probabilistically Triggered Arms},
  author = {Alihan Hüyük and Cem Tekin},
  journal= {arXiv preprint arXiv:1809.02707},
  year   = {2019}
}

备注

To appear in the Proceedings of the 22nd International Conference on Artificial Intelligence and Statistics (AISTATS) 2019