关于随机_bandit_的 KL-UCB+ 策略的注记
机器学习
2019-03-21 v2 机器学习
摘要
本注记考虑随机 K 臂 bandit 问题的经典设定。在该问题中,已知 KL-UCB 策略达到渐近最优后悔界,且 KL-UCB+ 策略在经验上优于 KL-UCB 策略,尽管原始形式 KL-UCB+ 策略的后悔界此前未知。本注记表明,可借助与其他已知策略分析相同的技术,给出 KL-UCB+ 策略渐近最优性的一个简单证明。
引用
@article{arxiv.1903.07839,
title = {A Note on KL-UCB+ Policy for the Stochastic Bandit},
author = {Junya Honda},
journal= {arXiv preprint arXiv:1903.07839},
year = {2019}
}
备注
6 pages, corrected typos