通过高斯逼近实现$\tilde{\mathcal{O}}(1/N)$最优性间隙的游荡匪徒
最优化与控制
2025-05-27 v2 机器学习
概率论
摘要
我们研究具有个同质臂的有限时域游荡多臂匪徒(RMAB)问题。先前工作表明,当RMAB满足非退化条件时,基于流体逼近(捕捉系统均值动力学)的线性规划(LP)策略可实现指数小的最优性间隙。然而,RMAB通常是退化的,此时基于LP的策略可能导致每臂的最优性间隙。本文提出了一种新颖的基于随机规划(SP)的策略,在唯一性假设下,对于退化RMAB实现了的最优性间隙。我们的方法基于构建一个高斯随机系统,该系统不仅捕捉RMAB动力学的均值,还捕捉其方差,从而比流体逼近更精确。然后我们求解该系统的随机规划以获得策略。这是首个为退化RMAB建立最优性间隙的结果。
引用
@article{arxiv.2410.15003,
title = {Achieving $\tilde{\mathcal{O}}(1/N)$ Optimality Gap in Restless Bandits through Gaussian Approximation},
author = {Chen Yan and Weina Wang and Lei Ying},
journal= {arXiv preprint arXiv:2410.15003},
year = {2025}
}
备注
53 pages, 5 figures. We weaken the assumption underlying our main theoretical result and also prove a more general version that holds without this condition