中文

势 Mean Field Games 中最优策略的学习:平滑策略迭代算法

最优化与控制 2023-04-18 v2

摘要

我们引入两种平滑策略迭代算法(\textbf{SPI})作为二阶势 Mean Field Games (MFGs) 中学习策略和计算纳什均衡的方法。若 MFG 系统中的耦合项满足 Lasry Lions 单调性条件,则证明了全局收敛性。对于可能具有多个解的系统,证明了到稳定解的局部收敛性。收敛分析显示了 \textbf{SPI} 与虚构博弈 (Fictitious Play) 算法之间的密切联系,后者在 MFG 文献中已被广泛研究。给出了基于有限差分格式的数值模拟结果以补充理论分析。

关键词

引用

@article{arxiv.2212.04791,
  title  = {Learning optimal policies in potential Mean Field Games: Smoothed Policy Iteration algorithms},
  author = {Qing Tang and Jiahao Song},
  journal= {arXiv preprint arXiv:2212.04791},
  year   = {2023}
}