English

Learning Approximate Nash Equilibria in Cooperative Multi-Agent Reinforcement Learning via Mean-Field Subsampling

Multiagent Systems 2026-05-12 v2 Artificial Intelligence Machine Learning Systems and Control Systems and Control Optimization and Control

Abstract

Many large-scale platforms and networked control systems have a centralized decision maker interacting with a massive population of agents under strict observability constraints. Motivated by such applications, we study a cooperative Markov game with a global agent and nn homogeneous local agents in a communication-constrained regime, where the global agent only observes a subset of kk local agent states per time step. We propose an alternating learning framework (ALTERNATING-MARL)(\texttt{ALTERNATING-MARL}), where the global agent performs subsampled mean-field QQ-learning against a fixed local policy, and local agents update by optimizing in an induced MDP. We prove that these approximate best-response dynamics converge to an O~(1/k)\widetilde{O}(1/\sqrt{k})-approximate Nash Equilibrium, while separating the sample complexities between the joint state and action spaces. Finally, we validate our results in numerical simulations for multi-robot control.

Keywords

Cite

@article{arxiv.2603.03759,
  title  = {Learning Approximate Nash Equilibria in Cooperative Multi-Agent Reinforcement Learning via Mean-Field Subsampling},
  author = {Emile Anand and Ishani Karmarkar},
  journal= {arXiv preprint arXiv:2603.03759},
  year   = {2026}
}

Comments

57 pages, 10 figures, 4 tables

R2 v1 2026-07-01T11:02:31.428Z