折扣成本离散时间系统的策略迭代:稳定性与近最优性保证
最优化与控制
2024-03-29 v1
摘要
给定折扣成本,我们研究输入由策略迭代(PI)生成的确定性离散时间系统。我们提供了新颖的近最优性和稳定性性质,同时允许初始策略非稳定。即,我们首先给出了由 PI 生成的值函数与最优值函数之间差异的新颖界限,对于所考虑的系统类别,这些界限通常不如动态规划文献中遇到的界限保守。然后,我们证明了在温和条件下,与 PI 生成的策略构成闭环的系统在有限(且已知)次数的迭代后是稳定的。
引用
@article{arxiv.2403.19007,
title = {Policy iteration for discrete-time systems with discounted costs: stability and near-optimality guarantees},
author = {Jonathan de Brusse and Mathieu Granzotto and Romain Postoyan and Dragan Nešić},
journal= {arXiv preprint arXiv:2403.19007},
year = {2024}
}