中文

基于深度强化学习的Population-aware在线镜像梯度均值场游戏

机器学习 2025-09-04 v1 多智能体系统 机器人学 系统与控制 系统与控制

摘要

均值场游戏(MFGs)为研究大规模多智能体系统提供了强大的框架。然而,在初始分布未知或受到common noise影响时学习Nash均衡仍是一个具有挑战性的问题。本文引入了一个高效的深度强化学习(DRL)算法,旨在不依赖于平均或历史抽样即可实现population-dependent Nash均衡,灵感来自Munchausen RL和Online Mirror Descent。得到的策略适用于各种初始分布和various sources of common noise。通过对七个标准例子的numerical实验,我们展示了该算法在population-dependent策略方面具有优于现有算法的收敛性能,特别是针对population-dependent策略的DRL版本的Fictitious Play。面对common noise的性能进一步突显了该方法的robust性和适应性。

关键词

引用

@article{arxiv.2509.03030,
  title  = {Population-aware Online Mirror Descent for Mean-Field Games with Common Noise by Deep Reinforcement Learning},
  author = {Zida Wu and Mathieu Lauriere and Matthieu Geist and Olivier Pietquin and Ankur Mehta},
  journal= {arXiv preprint arXiv:2509.03030},
  year   = {2025}
}

备注

2025 IEEE 64rd Conference on Decision and Control (CDC)