English

Kiefer Wolfowitz Algorithm is Asymptotically Optimal for a Class of Non-Stationary Bandit Problems

Machine Learning 2017-03-09 v2 Machine Learning

Abstract

We consider the problem of designing an allocation rule or an "online learning algorithm" for a class of bandit problems in which the set of control actions available at each time ss is a convex, compact subset of Rd\mathbb{R}^d. Upon choosing an action xx at time ss, the algorithm obtains a noisy value of the unknown and time-varying function fsf_s evaluated at xx. The "regret" of an algorithm is the gap between its expected reward, and the reward earned by a strategy which has the knowledge of the function fsf_s at each time ss and hence chooses the action xsx_s that maximizes fsf_s. For this non-stationary bandit problem set-up, we consider two variants of the Kiefer Wolfowitz (KW) algorithm i) KW with fixed step-size β\beta, and ii) KW with sliding window of length LL. We show that if the number of times that the function fsf_s varies during time TT is o(T)o(T), and if the learning rates of the proposed algorithms are chosen "optimally", then the regret of the proposed algorithms is o(T)o(T), and hence the algorithms are asymptotically efficient.

Keywords

Cite

@article{arxiv.1702.08000,
  title  = {Kiefer Wolfowitz Algorithm is Asymptotically Optimal for a Class of Non-Stationary Bandit Problems},
  author = {Rahul Singh and Taposh Banerjee},
  journal= {arXiv preprint arXiv:1702.08000},
  year   = {2017}
}
R2 v1 2026-06-22T18:28:39.128Z