Kiefer Wolfowitz Algorithm is Asymptotically Optimal for a Class of Non-Stationary Bandit Problems
Abstract
We consider the problem of designing an allocation rule or an "online learning algorithm" for a class of bandit problems in which the set of control actions available at each time is a convex, compact subset of . Upon choosing an action at time , the algorithm obtains a noisy value of the unknown and time-varying function evaluated at . The "regret" of an algorithm is the gap between its expected reward, and the reward earned by a strategy which has the knowledge of the function at each time and hence chooses the action that maximizes . For this non-stationary bandit problem set-up, we consider two variants of the Kiefer Wolfowitz (KW) algorithm i) KW with fixed step-size , and ii) KW with sliding window of length . We show that if the number of times that the function varies during time is , and if the learning rates of the proposed algorithms are chosen "optimally", then the regret of the proposed algorithms is , and hence the algorithms are asymptotically efficient.
Cite
@article{arxiv.1702.08000,
title = {Kiefer Wolfowitz Algorithm is Asymptotically Optimal for a Class of Non-Stationary Bandit Problems},
author = {Rahul Singh and Taposh Banerjee},
journal= {arXiv preprint arXiv:1702.08000},
year = {2017}
}