中文

Copeland对决老虎机

机器学习 2015-06-02 v1

摘要

研究了一类对决老虎机问题,其中可能不存在Condorcet winner。提出了两种算法,转而寻求最小化相对于Copeland winner的后悔值,与Condorcet winner不同,Copeland winner保证存在。第一种算法Copeland Confidence Bound (CCB)针对少量臂设计,而第二种算法Scalable Copeland Bandits (SCB)在大规模问题上表现更好。我们提供了限制CCB和SCB累积后悔值的理论结果,均大幅改进了现有结果。现有结果要么提供形式为O(KlogT)O(K \log T)的界限但需要限制性假设,要么提供形式为O(K2logT)O(K^2 \log T)的界限而不需要此类假设。我们的结果兼得二者优势:无需限制性假设即可得到O(KlogT)O(K \log T)界限。

关键词

引用

@article{arxiv.1506.00312,
  title  = {Copeland Dueling Bandits},
  author = {Masrour Zoghi and Zohar Karnin and Shimon Whiteson and Maarten de Rijke},
  journal= {arXiv preprint arXiv:1506.00312},
  year   = {2015}
}

备注

33 pages, 8 figures