中文

兼具对决与拉取操作的阈值赌博机问题

机器学习 2020-06-16 v2 机器学习

摘要

阈值赌博机问题(TBP)旨在找出均值奖励大于给定阈值的臂集合。我们考虑TBP的一种新设置,其中除拉取臂外,还可以对决两个臂并获得均值较大的那个臂。在我们来自众包的激励应用中,对决两个臂比直接拉取更具成本效益且更省时。我们将此问题称为带对决选择的TBP(TBP-DC)。本文提供了一种称为Rank-Search(RS)的算法,通过排序与二分搜索交替来解决TBP-DC。我们证明了RS的理论保证,并给出下界以表明其最优性。实验表明RS优于仅使用拉取或对决的先前基线算法。

关键词

引用

@article{arxiv.1910.06368,
  title  = {Thresholding Bandit Problem with Both Duels and Pulls},
  author = {Yichong Xu and Xi Chen and Aarti Singh and Artur Dubrawski},
  journal= {arXiv preprint arXiv:1910.06368},
  year   = {2020}
}

备注

15 pages, 8 figures; The 23rd International Conference on Artificial Intelligence and Statistics