中文

随机多臂赌博机问题中的遗憾下界与扩展 Upper Confidence Bounds 策略

机器学习 2011-12-19 v1

摘要

本文致力于研究经典随机多臂赌博机模型中的遗憾下界。Lai 和 Robbins 的一个著名结果(随后被 Burnetas 和 Katehakis 推广)确立了对所有一致策略存在对数界。我们放宽了一致性的概念,并给出了该对数界的一个推广。我们还证明了在 Hannan 一致性的一般情形下不存在对数界。为了获得这些结果,我们研究了流行的 Upper Confidence Bounds (ucb) 策略的变体。作为副产品,我们证明了不可能设计出一种自适应策略,能够通过利用环境的性质来从两种算法中选择最优者。

关键词

引用

@article{arxiv.1112.3827,
  title  = {Regret lower bounds and extended Upper Confidence Bounds policies in stochastic multi-armed bandit problem},
  author = {Antoine Salomon and Jean-Yves Audibert and Issam El Alaoui},
  journal= {arXiv preprint arXiv:1112.3827},
  year   = {2011}
}