中文

对数线性策略下自然策略梯度方法的线性收敛性

机器学习 2023-02-22 v3 人工智能 最优化与控制

摘要

我们考虑无限 horizon 折扣马尔可夫决策过程,并研究自然策略梯度(NPG)与 Q-NPG 方法在对数线性策略类下的收敛速率。利用兼容函数近似框架,两种采用对数线性策略的方法均可写为策略镜像下降(PMD)方法的不精确版本。我们表明,两种方法在使用简单的、非自适应的几何递增步长时,均达到线性收敛速率与 O~(1/ϵ2)\tilde{\mathcal{O}}(1/\epsilon^2) 的样本复杂度,而无需借助熵或其他强凸正则化。最后,作为副产品,我们得到了两种方法在任意常数步长下的次线性收敛速率。

关键词

引用

@article{arxiv.2210.01400,
  title  = {Linear Convergence of Natural Policy Gradient Methods with Log-Linear Policies},
  author = {Rui Yuan and Simon S. Du and Robert M. Gower and Alessandro Lazaric and Lin Xiao},
  journal= {arXiv preprint arXiv:2210.01400},
  year   = {2023}
}

备注

This version adds a table of comparison for the literature review. The paper is published as a conference paper at ICLR 2023