English

Breaking the Total Variance Barrier: Sharp Sample Complexity for Linear Heteroscedastic Bandits with Fixed Action Set

Machine Learning 2026-07-26 v1 Machine Learning

Abstract

Recent years have witnessed increasing interests in tackling heteroscedastic noise in bandits and reinforcement learning. In these works, the cumulative variance of the noise Λ=t=1Tσt2\Lambda = \sum_{t=1}^T \sigma_t^2, where σt2\sigma_t^2 is the variance of the noise at round tt, is used to characterize the statistical complexity of the problem, yielding \emph{simple regret} bounds of order O~(dΛ/T2)\tilde{\cal{O}}(d \sqrt{\Lambda / T^2}) for dd-dimensional linear bandits with heteroscedastic noise. However, with a closer look, Λ\Lambda remains the same order even if the noise is close to zero at half of the rounds, which indicates that the Λ\Lambda-dependence is not optimal. In this paper, we revisit the stochastic linear bandit problem with heteroscedastic noise, where the action set is prefixed throughout the learning process. We propose a novel variance-adaptive algorithm \texttt{VAEE} (Variance-Aware Exploration with Elimination) for large action set, which actively explores actions that maximizes the information gain among a candidate set of actions that are not eliminated. With the active-exploration strategy, we show that \texttt{VAEE} achieves a \emph{simple regret} with a nearly \emph{harmonic-mean} dependent rate. For finitely many actions, we propose a variance-aware variant of G-optimal design based exploration, which achieves a simple regret with sharper dependence on dd. We also establish a nearly matching lower bound for the fixed action set setting indicating that \emph{harmonic-mean} dependent rate is unavoidable. To the best of our knowledge, this is the first work that breaks the Λ\sqrt{\Lambda} barrier for stochastic linear bandits with heteroscedastic noise.

Cite

@article{arxiv.2607.23679,
  title  = {Breaking the Total Variance Barrier: Sharp Sample Complexity for Linear Heteroscedastic Bandits with Fixed Action Set},
  author = {Heyang Zhao and Tianyuan Jin and Weixin Wang and Vincent Y. F. Tan and Pan Xu and Quanquan Gu},
  journal= {arXiv preprint arXiv:2607.23679},
  year   = {2026}
}