Learning Approximate Nash Equilibria in Cooperative Multi-Agent Reinforcement Learning via Mean-Field Subsampling
Abstract
Many large-scale platforms and networked control systems have a centralized decision maker interacting with a massive population of agents under strict observability constraints. Motivated by such applications, we study a cooperative Markov game with a global agent and homogeneous local agents in a communication-constrained regime, where the global agent only observes a subset of local agent states per time step. We propose an alternating learning framework , where the global agent performs subsampled mean-field -learning against a fixed local policy, and local agents update by optimizing in an induced MDP. We prove that these approximate best-response dynamics converge to an -approximate Nash Equilibrium, while separating the sample complexities between the joint state and action spaces. Finally, we validate our results in numerical simulations for multi-robot control.
Cite
@article{arxiv.2603.03759,
title = {Learning Approximate Nash Equilibria in Cooperative Multi-Agent Reinforcement Learning via Mean-Field Subsampling},
author = {Emile Anand and Ishani Karmarkar},
journal= {arXiv preprint arXiv:2603.03759},
year = {2026}
}
Comments
57 pages, 10 figures, 4 tables