English

Minimax Rate-Optimal Algorithms for High-Dimensional Stochastic Linear Bandits

Statistics Theory 2025-05-26 v1 Machine Learning Statistics Theory

Abstract

We study the stochastic linear bandit problem with multiple arms over TT rounds, where the covariate dimension dd may exceed TT, but each arm-specific parameter vector is ss-sparse. We begin by analyzing the sequential estimation problem in the single-arm setting, focusing on cumulative mean-squared error. We show that Lasso estimators are provably suboptimal in the sequential setting, exhibiting suboptimal dependence on dd and TT, whereas thresholded Lasso estimators -- obtained by applying least squares to the support selected by thresholding an initial Lasso estimator -- achieve the minimax rate. Building on these insights, we consider the full linear contextual bandit problem and propose a three-stage arm selection algorithm that uses thresholded Lasso as the main estimation method. We derive an upper bound on the cumulative regret of order s(logs)(logd+logT)s(\log s)(\log d + \log T), and establish a matching lower bound up to a logs\log s factor, thereby characterizing the minimax regret rate up to a logarithmic term in ss. Moreover, when a short initial period is excluded from the regret, the proposed algorithm achieves exact minimax optimality.

Keywords

Cite

@article{arxiv.2505.17400,
  title  = {Minimax Rate-Optimal Algorithms for High-Dimensional Stochastic Linear Bandits},
  author = {Jingyu Liu and Yanglei Song},
  journal= {arXiv preprint arXiv:2505.17400},
  year   = {2025}
}
R2 v1 2026-07-01T02:33:00.054Z