对数线性策略下自然策略梯度方法的线性收敛性
机器学习
2023-02-22 v3 人工智能
最优化与控制
摘要
我们考虑无限 horizon 折扣马尔可夫决策过程,并研究自然策略梯度(NPG)与 Q-NPG 方法在对数线性策略类下的收敛速率。利用兼容函数近似框架,两种采用对数线性策略的方法均可写为策略镜像下降(PMD)方法的不精确版本。我们表明,两种方法在使用简单的、非自适应的几何递增步长时,均达到线性收敛速率与 的样本复杂度,而无需借助熵或其他强凸正则化。最后,作为副产品,我们得到了两种方法在任意常数步长下的次线性收敛速率。
引用
@article{arxiv.2210.01400,
title = {Linear Convergence of Natural Policy Gradient Methods with Log-Linear Policies},
author = {Rui Yuan and Simon S. Du and Robert M. Gower and Alessandro Lazaric and Lin Xiao},
journal= {arXiv preprint arXiv:2210.01400},
year = {2023}
}
备注
This version adds a table of comparison for the literature review. The paper is published as a conference paper at ICLR 2023