English

On The Complexity of Best-Arm Identification in Non-Stationary Linear Bandits

Machine Learning 2026-03-12 v1 Machine Learning

Abstract

We study the fixed-budget best-arm identification (BAI) problem in non-stationary linear bandits. Concretely, given a fixed time budget TNT\in \mathbb{N}, finite arm set XRd\mathcal{X} \subset \mathbb{R}^d, and a potentially adversarial sequence of unknown parameters {θt}t=1T\lbrace \theta_t\rbrace_{t=1}^{T} (hence non-stationary), a learner aims to identify the arm with the largest cumulative reward x=argmaxxXxt=1Tθtx_* = \arg\max_{x \in \mathcal{X}} x^\top\sum_{t=1}^T \theta_t with high probability. In this setting, it is well-known that uniformly sampling arms from the G-optimal design yields a minimax-optimal error probability of exp(Θ(T/HG))\exp\left(-\Theta\left(T / H_{G}\right)\right), where HGH_{G} scales proportionally with the dimension dd. However, this notion of complexity is overly pessimistic, as it is derived from a lower bound in which the arm set consists only of the standard basis vectors, thus masking any potential advantages arising from arm sets with richer geometric structure. To address this, we establish an arm-set-dependent lower bound that, in contrast, holds for any arm set. Motivated by the ideas underlying our lower bound, we propose the Adjacent-optimal design, a specialization of the well-known XY\mathcal{X}\mathcal{Y}-optimal design, and develop the Adjacent-BAI\textsf{Adjacent-BAI} algorithm. We prove that the error probability of Adjacent-BAI\textsf{Adjacent-BAI} matches our lower bound up to constants, verifying the tightness of our lower bound, and establishing the arm-set-dependent complexity of this setting.

Keywords

Cite

@article{arxiv.2603.10346,
  title  = {On The Complexity of Best-Arm Identification in Non-Stationary Linear Bandits},
  author = {Leo Maynard-Zhang and Zhihan Xiong and Kevin Jamieson and Maryam Fazel},
  journal= {arXiv preprint arXiv:2603.10346},
  year   = {2026}
}