English

Asymptotic Convergence of Thompson Sampling

Machine Learning 2020-11-10 v1 Statistics Theory Statistics Theory

Abstract

Thompson sampling has been shown to be an effective policy across a variety of online learning tasks. Many works have analyzed the finite time performance of Thompson sampling, and proved that it achieves a sub-linear regret under a broad range of probabilistic settings. However its asymptotic behavior remains mostly underexplored. In this paper, we prove an asymptotic convergence result for Thompson sampling under the assumption of a sub-linear Bayesian regret, and show that the actions of a Thompson sampling agent provide a strongly consistent estimator of the optimal action. Our results rely on the martingale structure inherent in Thompson sampling.

Keywords

Cite

@article{arxiv.2011.03917,
  title  = {Asymptotic Convergence of Thompson Sampling},
  author = {Cem Kalkanli and Ayfer Ozgur},
  journal= {arXiv preprint arXiv:2011.03917},
  year   = {2020}
}
R2 v1 2026-06-23T19:59:20.691Z