Whittle index based Q-learning for restless bandits with average reward
Machine Learning
2021-09-22 v3 Optimization and Control
Machine Learning
Abstract
A novel reinforcement learning algorithm is introduced for multiarmed restless bandits with average reward, using the paradigms of Q-learning and Whittle index. Specifically, we leverage the structure of the Whittle index policy to reduce the search space of Q-learning, resulting in major computational gains. Rigorous convergence analysis is provided, supported by numerical experiments. The numerical experiments show excellent empirical performance of the proposed scheme.
Keywords
Cite
@article{arxiv.2004.14427,
title = {Whittle index based Q-learning for restless bandits with average reward},
author = {Konstantin E. Avrachenkov and Vivek S. Borkar},
journal= {arXiv preprint arXiv:2004.14427},
year = {2021}
}