English

Multi-Agent Guided Policy Search for Non-Cooperative Dynamic Games

Systems and Control 2026-02-13 v4 Systems and Control

Abstract

Multi-agent reinforcement learning (MARL) optimizes strategic interactions in non-cooperative dynamic games, where agents have misaligned objectives. However, data-driven methods such as multi-agent policy gradients (MA-PG) often suffer from instability and limit-cycle behaviors. Prior stabilization techniques typically rely on entropy-based exploration, which slows learning and increases variance. We propose a model-based approach that incorporates approximate priors into the reward function as regularization. In linear quadratic (LQ) games, we prove that such priors stabilize policy gradients and guarantee local exponential convergence to an approximate Nash equilibrium. We then extend this idea to infinite-horizon nonlinear games by introducing Multi-agent Guided Policy Search (MA-GPS), which constructs short-horizon local LQ approximations from trajectories of current policies to guide training. Experiments on nonlinear vehicle platooning and a six-player strategic basketball formation show that MA-GPS achieves faster convergence and more stable learning than existing MARL methods.

Keywords

Cite

@article{arxiv.2509.24226,
  title  = {Multi-Agent Guided Policy Search for Non-Cooperative Dynamic Games},
  author = {Jingqi Li and Gechen Qu and Jason J. Choi and Somayeh Sojoudi and Claire Tomlin},
  journal= {arXiv preprint arXiv:2509.24226},
  year   = {2026}
}

Comments

This paper has been accepted for presentation at the IEEE American Control Conference (ACC) 2026. We sincerely appreciate the reviewers' valuable and constructive feedback. The latest version of the manuscript incorporates their suggestions, including additional clarifications of theoretical assumptions, convergence guarantees, and experimental details

R2 v1 2026-07-01T06:03:27.163Z