基于深度强化学习的Population-aware在线镜像梯度均值场游戏
机器学习
2025-09-04 v1 多智能体系统
机器人学
系统与控制
系统与控制
摘要
均值场游戏(MFGs)为研究大规模多智能体系统提供了强大的框架。然而,在初始分布未知或受到common noise影响时学习Nash均衡仍是一个具有挑战性的问题。本文引入了一个高效的深度强化学习(DRL)算法,旨在不依赖于平均或历史抽样即可实现population-dependent Nash均衡,灵感来自Munchausen RL和Online Mirror Descent。得到的策略适用于各种初始分布和various sources of common noise。通过对七个标准例子的numerical实验,我们展示了该算法在population-dependent策略方面具有优于现有算法的收敛性能,特别是针对population-dependent策略的DRL版本的Fictitious Play。面对common noise的性能进一步突显了该方法的robust性和适应性。
引用
@article{arxiv.2509.03030,
title = {Population-aware Online Mirror Descent for Mean-Field Games with Common Noise by Deep Reinforcement Learning},
author = {Zida Wu and Mathieu Lauriere and Matthieu Geist and Olivier Pietquin and Ankur Mehta},
journal= {arXiv preprint arXiv:2509.03030},
year = {2025}
}
备注
2025 IEEE 64rd Conference on Decision and Control (CDC)