English

Chained Information-Theoretic bounds and Tight Regret Rate for Linear Bandit Problems

Machine Learning 2024-03-07 v1 Machine Learning

Abstract

This paper studies the Bayesian regret of a variant of the Thompson-Sampling algorithm for bandit problems. It builds upon the information-theoretic framework of [Russo and Van Roy, 2015] and, more specifically, on the rate-distortion analysis from [Dong and Van Roy, 2020], where they proved a bound with regret rate of O(dTlog(T))O(d\sqrt{T \log(T)}) for the dd-dimensional linear bandit setting. We focus on bandit problems with a metric action space and, using a chaining argument, we establish new bounds that depend on the metric entropy of the action space for a variant of Thompson-Sampling. Under suitable continuity assumption of the rewards, our bound offers a tight rate of O(dT)O(d\sqrt{T}) for dd-dimensional linear bandit problems.

Keywords

Cite

@article{arxiv.2403.03361,
  title  = {Chained Information-Theoretic bounds and Tight Regret Rate for Linear Bandit Problems},
  author = {Amaury Gouverneur and Borja Rodríguez-Gálvez and Tobias J. Oechtering and Mikael Skoglund},
  journal= {arXiv preprint arXiv:2403.03361},
  year   = {2024}
}

Comments

15 pages: 8 of main text and 7 of appendices

R2 v1 2026-06-28T15:10:26.712Z