具有最优聚合遗憾的阈值赌博机
机器学习
2019-05-28 v1 机器学习
摘要
我们考虑阈值赌博机问题,其目标是在 次试验的固定预算下,找出平均奖励高于给定阈值 的臂。我们引入 LSA,一种新颖、简单且任意时刻可用的算法,旨在最小化聚合遗憾(或误分类臂的期望数量)。我们证明我们的算法在实例意义下是渐近最优的。我们还提供了全面的实证结果,以证明该算法在多种不同场景下相对于现有算法的优越性能。
引用
@article{arxiv.1905.11046,
title = {Thresholding Bandit with Optimal Aggregate Regret},
author = {Chao Tao and Saùl Blanco and Jian Peng and Yuan Zhou},
journal= {arXiv preprint arXiv:1905.11046},
year = {2019}
}