English

Prior Diffusiveness and Regret in the Linear-Gaussian Bandit

Machine Learning 2026-01-06 v1

Abstract

We prove that Thompson sampling exhibits O~(σdT+drTr(Σ0))\tilde{O}(\sigma d \sqrt{T} + d r \sqrt{\mathrm{Tr}(\Sigma_0)}) Bayesian regret in the linear-Gaussian bandit with a N(μ0,Σ0)\mathcal{N}(\mu_0, \Sigma_0) prior distribution on the coefficients, where dd is the dimension, TT is the time horizon, rr is the maximum 2\ell_2 norm of the actions, and σ2\sigma^2 is the noise variance. In contrast to existing regret bounds, this shows that to within logarithmic factors, the prior-dependent ``burn-in'' term drTr(Σ0)d r \sqrt{\mathrm{Tr}(\Sigma_0)} decouples additively from the minimax (long run) regret σdT\sigma d \sqrt{T}. Previous regret bounds exhibit a multiplicative dependence on these terms. We establish these results via a new ``elliptical potential'' lemma, and also provide a lower bound indicating that the burn-in term is unavoidable.

Cite

@article{arxiv.2601.02022,
  title  = {Prior Diffusiveness and Regret in the Linear-Gaussian Bandit},
  author = {Yifan Zhu and John C. Duchi and Benjamin Van Roy},
  journal= {arXiv preprint arXiv:2601.02022},
  year   = {2026}
}
R2 v1 2026-07-01T08:50:43.829Z