Prior Diffusiveness and Regret in the Linear-Gaussian Bandit
Machine Learning
2026-01-06 v1
Abstract
We prove that Thompson sampling exhibits Bayesian regret in the linear-Gaussian bandit with a prior distribution on the coefficients, where is the dimension, is the time horizon, is the maximum norm of the actions, and is the noise variance. In contrast to existing regret bounds, this shows that to within logarithmic factors, the prior-dependent ``burn-in'' term decouples additively from the minimax (long run) regret . Previous regret bounds exhibit a multiplicative dependence on these terms. We establish these results via a new ``elliptical potential'' lemma, and also provide a lower bound indicating that the burn-in term is unavoidable.
Cite
@article{arxiv.2601.02022,
title = {Prior Diffusiveness and Regret in the Linear-Gaussian Bandit},
author = {Yifan Zhu and John C. Duchi and Benjamin Van Roy},
journal= {arXiv preprint arXiv:2601.02022},
year = {2026}
}