面向支持感知 CVaR 多臂老虎机的优化 Thompson Sampling 策略
机器学习
2022-03-22 v3
摘要
本文研究了一种多臂老虎机问题,其中每个臂的质量由奖励分布在某一水平 alpha 下的条件风险价值(Conditional Value at Risk, CVaR)来衡量。现有该设定下的工作主要集中于上置信界(Upper Confidence Bound)算法,我们则针对有界奖励的 CVaR 老虎机引入了一种新的 Thompson Sampling 方法,该方法具有足够的灵活性,可解决基于物理资源的多种问题。基于 Riou & Honda (2020) 的近期工作,我们针对连续有界奖励引入了 B-CVTS,并针对多项分布引入了 M-CVTS。在理论方面,我们对其分析进行了非平凡的扩展,从而能够从理论上界定其 CVaR 遗憾最小化性能。引人注目的是,我们的结果表明这些策略是首批被证明可在 CVaR 老虎机中实现渐近最优性的策略,与该设定下相应的渐近下界相匹配。此外,我们在模拟农业用例的现实环境以及多种合成示例中实证说明了 Thompson Sampling 方法的优势。
引用
@article{arxiv.2012.05754,
title = {Optimal Thompson Sampling strategies for support-aware CVaR bandits},
author = {Dorian Baudry and Romain Gautron and Emilie Kaufmann and Odalric-Ambryn Maillard},
journal= {arXiv preprint arXiv:2012.05754},
year = {2022}
}
备注
Presented at the Thirty-eighth International Conference on Machine Learning (ICML 2021). In this version we refine Lemma 2 and correct its proof (does not change the main theorems)