中文

分布式 Thompson Sampling

人工智能 2021-09-10 v2 机器学习

摘要

我们研究具有 M 个智能体与 K 个臂的协作式多智能体多臂老虎机问题。智能体的目标是最小化累积后悔。我们在分布式设定下改造了传统的 Thompson Sampling 算法。然而,借助智能体的通信能力,我们注意到通信可进一步降低分布式 Thompson Sampling 方法的后悔上界。为进一步提升分布式 Thompson Sampling 的性能,我们提出一种基于消除的分布式 Thompson Sampling 算法,允许智能体协作学习。我们在伯努利奖励下分析了该算法,并推导了依赖于问题的累积后悔上界。

关键词

引用

@article{arxiv.2012.01789,
  title  = {Distributed Thompson Sampling},
  author = {Jing Dong and Tan Li and Shaolei Ren and Linqi Song},
  journal= {arXiv preprint arXiv:2012.01789},
  year   = {2021}
}

备注

The paper is not finished and will not be updated