中文

关于随机_bandit_的 KL-UCB+ 策略的注记

机器学习 2019-03-21 v2 机器学习

摘要

本注记考虑随机 K 臂 bandit 问题的经典设定。在该问题中,已知 KL-UCB 策略达到渐近最优后悔界,且 KL-UCB+ 策略在经验上优于 KL-UCB 策略,尽管原始形式 KL-UCB+ 策略的后悔界此前未知。本注记表明,可借助与其他已知策略分析相同的技术,给出 KL-UCB+ 策略渐近最优性的一个简单证明。

关键词

引用

@article{arxiv.1903.07839,
  title  = {A Note on KL-UCB+ Policy for the Stochastic Bandit},
  author = {Junya Honda},
  journal= {arXiv preprint arXiv:1903.07839},
  year   = {2019}
}

备注

6 pages, corrected typos