NeurWIN:基于深度强化学习的 restless _bandit 神经 Whittle 指数网络
机器学习
2022-01-21 v2 机器学习
摘要
Whittle 指数策略是为 notorious 难解的 restless bandit 问题获得渐近最优解的有力工具。然而,对于许多具有复杂转移核的实际 restless bandit,寻找 Whittle 指数仍然是一个困难问题。本文提出 NeurWIN,一种神经 Whittle 指数网络,通过利用 Whittle 指数的数学性质来为任意 restless bandit 学习 Whittle 指数。我们表明,产生 Whittle 指数的神经网络也是为一组合适的马尔可夫决策问题产生最优控制的网络。这一性质激励使用深度强化学习来训练 NeurWIN。我们通过评估其在三个近期研究的 restless bandit 问题上的性能来展示 NeurWIN 的效用。我们的实验结果表明,NeurWIN 的性能显著优于其他 RL 算法。
引用
@article{arxiv.2110.02128,
title = {NeurWIN: Neural Whittle Index Network For Restless Bandits Via Deep RL},
author = {Khaled Nakhleh and Santosh Ganji and Ping-Chun Hsieh and I-Hong Hou and Srinivas Shakkottai},
journal= {arXiv preprint arXiv:2110.02128},
year = {2022}
}
备注
Accepted for publication in NeurIPS 2021