面向上下文感知_bandit 的双线性 Thompson 采样
机器学习
2020-10-20 v1 人工智能
摘要
在本文中,我们分析并扩展了一种称为上下文感知_bandit(Context-Attentive Bandit)的在线学习框架,该框架受多种实际应用驱动,从医学诊断到对话系统,其中由于观测成本,在每次迭代中只能观测到潜在大量上下文变量中的一小部分子集;然而,智能体可自由选择观测哪些变量。我们推导了一种新颖算法,称为上下文感知 Thompson 采样(CATS),它构建于线性 Thompson 采样方法之上,并将其适配到上下文感知_bandit 设定中。我们提供了理论后悔分析以及广泛的实证评估,在多种真实数据集上展示了所提方法相对于若干基线方法的优势。
引用
@article{arxiv.2010.09473,
title = {Double-Linear Thompson Sampling for Context-Attentive Bandits},
author = {Djallel Bouneffouf and Raphaël Féraud and Sohini Upadhyay and Yasaman Khazaeni and Irina Rish},
journal= {arXiv preprint arXiv:2010.09473},
year = {2020}
}
备注
arXiv admin note: text overlap with arXiv:1906.09384