An Asymptotically Optimal Index Policy for Finite-Horizon Restless Bandits
Optimization and Control
2017-07-04 v1
Abstract
We consider restless multi-armed bandit (RMAB) with a finite horizon and multiple pulls per period. Leveraging the Lagrangian relaxation, we approximate the problem with a collection of single arm problems. We then propose an index-based policy that uses optimal solutions of the single arm problems to index individual arms, and offer a proof that it is asymptotically optimal as the number of arms tends to infinity. We also use simulation to show that this index-based policy performs better than the state-of-art heuristics in various problem settings.
Keywords
Cite
@article{arxiv.1707.00205,
title = {An Asymptotically Optimal Index Policy for Finite-Horizon Restless Bandits},
author = {Weici Hu and Peter Frazier},
journal= {arXiv preprint arXiv:1707.00205},
year = {2017}
}