分布式 Thompson Sampling
人工智能
2021-09-10 v2 机器学习
摘要
我们研究具有 M 个智能体与 K 个臂的协作式多智能体多臂老虎机问题。智能体的目标是最小化累积后悔。我们在分布式设定下改造了传统的 Thompson Sampling 算法。然而,借助智能体的通信能力,我们注意到通信可进一步降低分布式 Thompson Sampling 方法的后悔上界。为进一步提升分布式 Thompson Sampling 的性能,我们提出一种基于消除的分布式 Thompson Sampling 算法,允许智能体协作学习。我们在伯努利奖励下分析了该算法,并推导了依赖于问题的累积后悔上界。
引用
@article{arxiv.2012.01789,
title = {Distributed Thompson Sampling},
author = {Jing Dong and Tan Li and Shaolei Ren and Linqi Song},
journal= {arXiv preprint arXiv:2012.01789},
year = {2021}
}
备注
The paper is not finished and will not be updated