基于神经网络的策略迭代算法及其在区域上随机博弈的全局 $H^2$-超线性收敛
数值分析
2020-02-14 v3 机器学习
数值分析
最优化与控制
摘要
本工作中,我们提出一类求解半线性Hamilton-Jacobi-Bellman-Isaacs(HJBI)边值问题的数值格式,该类问题自然源于带受控漂移扩散过程的离出时间问题。我们利用策略迭代将半线性问题化为一列线性Dirichlet问题,随后以多层前馈神经网络拟设逼近。我们确立了数值解在 -范数下全局收敛,并进一步通过将算法解释为HJBI方程的不精确Newton迭代,证明该收敛为超线性。此外,我们从数值值函数构造最优反馈控制并推导其收敛性。数值格式与收敛结果随后推广至对应于带斜边界反射受控扩散过程的HJBI边值问题。我们给出随机Zermelo导航问题的数值实验,以阐释理论结果并展示方法的有效性。
引用
@article{arxiv.1906.02304,
title = {A neural network based policy iteration algorithm with global $H^2$-superlinear convergence for stochastic games on domains},
author = {Kazufumi Ito and Christoph Reisinger and Yufei Zhang},
journal= {arXiv preprint arXiv:1906.02304},
year = {2020}
}
备注
Additional numerical experiments have been included (on Pages 27-31) to show the proposed algorithm achieves a more stable and more rapid convergence than the existing neural network based methods within similar computational time