线性二次调节器中策略梯度估计器方差的分析
机器学习
2019-10-04 v1 机器学习
摘要
我们研究了在具有连续状态与动作空间、线性动力学、二次代价和高斯噪声的环境中,REINFORCE 策略梯度估计器的方差。这些简单环境使我们能够根据环境和噪声参数推导估计器方差的界。我们将界的预测与仿真实验中的经验方差进行了比较。
引用
@article{arxiv.1910.01249,
title = {Analyzing the Variance of Policy Gradient Estimators for the Linear-Quadratic Regulator},
author = {James A. Preiss and Sébastien M. R. Arnold and Chen-Yu Wei and Marius Kloft},
journal= {arXiv preprint arXiv:1910.01249},
year = {2019}
}
备注
Accepted at NeurIPS 2019 Workshop on Optimization Foundations for Reinforcement Learning. 7 pages + 6 pages appendix