中文

面向通用值函数逼近的强化学习随机化探索

机器学习 2021-10-27 v2 机器学习

摘要

我们提出了一种无模型强化学习算法,其灵感来自流行的随机最小二乘值迭代(RLSVI)算法以及乐观原理。与现有基于上置信界(UCB)的方法(通常计算上难以处理)不同,我们的算法仅通过用精心选择的独立同分布标量噪声扰动训练数据来驱动探索。为了在不借助UCB式奖励的情况下获得乐观值函数估计,我们引入了乐观奖励采样过程。当值函数可由函数类 F\mathcal{F} 表示时,我们的算法实现了 O~(poly(dEH)T)\widetilde{O}(\mathrm{poly}(d_EH)\sqrt{T}) 的最坏情况后悔界,其中 TT 为经过的时间,HH 为规划时域,dEd_Eeluder dimension\textit{eluder dimension}(消隐维数) of F\mathcal{F}。在线性设定下,我们的算法退化为LSVI-PHE(RLSVI的一种变体),其享有 O~(d3H3T)\widetilde{\mathcal{O}}(\sqrt{d^3H^3T}) 的后悔。我们以在已知困难探索任务上的实证评估补充了该理论。

关键词

引用

@article{arxiv.2106.07841,
  title  = {Randomized Exploration for Reinforcement Learning with General Value Function Approximation},
  author = {Haque Ishfaq and Qiwen Cui and Viet Nguyen and Alex Ayoub and Zhuoran Yang and Zhaoran Wang and Doina Precup and Lin F. Yang},
  journal= {arXiv preprint arXiv:2106.07841},
  year   = {2021}
}

备注

32 page, 5 figures, in Proceedings of the 38th International Conference on Machine Learning, PMLR 139, 2021