势 Mean Field Games 中最优策略的学习:平滑策略迭代算法
最优化与控制
2023-04-18 v2
摘要
我们引入两种平滑策略迭代算法(\textbf{SPI})作为二阶势 Mean Field Games (MFGs) 中学习策略和计算纳什均衡的方法。若 MFG 系统中的耦合项满足 Lasry Lions 单调性条件,则证明了全局收敛性。对于可能具有多个解的系统,证明了到稳定解的局部收敛性。收敛分析显示了 \textbf{SPI} 与虚构博弈 (Fictitious Play) 算法之间的密切联系,后者在 MFG 文献中已被广泛研究。给出了基于有限差分格式的数值模拟结果以补充理论分析。
引用
@article{arxiv.2212.04791,
title = {Learning optimal policies in potential Mean Field Games: Smoothed Policy Iteration algorithms},
author = {Qing Tang and Jiahao Song},
journal= {arXiv preprint arXiv:2212.04791},
year = {2023}
}