随机多臂赌博机问题中的遗憾下界与扩展 Upper Confidence Bounds 策略
机器学习
2011-12-19 v1
摘要
本文致力于研究经典随机多臂赌博机模型中的遗憾下界。Lai 和 Robbins 的一个著名结果(随后被 Burnetas 和 Katehakis 推广)确立了对所有一致策略存在对数界。我们放宽了一致性的概念,并给出了该对数界的一个推广。我们还证明了在 Hannan 一致性的一般情形下不存在对数界。为了获得这些结果,我们研究了流行的 Upper Confidence Bounds (ucb) 策略的变体。作为副产品,我们证明了不可能设计出一种自适应策略,能够通过利用环境的性质来从两种算法中选择最优者。
引用
@article{arxiv.1112.3827,
title = {Regret lower bounds and extended Upper Confidence Bounds policies in stochastic multi-armed bandit problem},
author = {Antoine Salomon and Jean-Yves Audibert and Issam El Alaoui},
journal= {arXiv preprint arXiv:1112.3827},
year = {2011}
}