中文

关于比 Lipschitz 更光滑的 Bandit 上的 Thompson Sampling

机器学习 2020-02-27 v2 机器学习

摘要

Thompson Sampling 是解决 bandit 与强化学习问题的成熟方法。然而其在连续臂 bandit 问题中的使用受到的关注相对较少。我们在包含真实函数的函数类与次指数观测噪声的弱条件下,首次给出了 Thompson Sampling 用于连续臂 bandit 的悔界。我们的界通过对 eluder 维数的分析得到,eluder 维数是最近提出的函数类复杂度度量,已被证明在次高斯观测噪声下界定更简单 bandit 问题的 Thompson Sampling 贝叶斯悔方面有用。我们推导了具有 Lipschitz 导数的函数类的 eluder 维数新界,并在多方面推广了先前的分析。

关键词

引用

@article{arxiv.2001.02323,
  title  = {On Thompson Sampling for Smoother-than-Lipschitz Bandits},
  author = {James A. Grant and David S. Leslie},
  journal= {arXiv preprint arXiv:2001.02323},
  year   = {2020}
}

备注

Accepted to AISTATS 2020. 26 pages, 2 figures