English

Wasserstein-Barycenter Consensus for Cooperative Multi-Agent Reinforcement Learning

Systems and Control 2025-06-19 v2 Systems and Control

Abstract

Cooperative multi-agent reinforcement learning (MARL) demands principled mechanisms to align heterogeneous policies while preserving the capacity for specialized behavior. We introduce a novel consensus framework that defines the team strategy as the entropic-regularized pp-Wasserstein barycenter of agents' joint state--action visitation measures. By augmenting each agent's policy objective with a soft penalty proportional to its Sinkhorn divergence from this barycenter, the proposed approach encourages coherent group behavior without enforcing rigid parameter sharing. We derive an algorithm that alternates between Sinkhorn-barycenter computation and policy-gradient updates, and we prove that, under standard Lipschitz and compactness assumptions, the maximal pairwise policy discrepancy contracts at a geometric rate. Empirical evaluation on a cooperative navigation case study demonstrates that our OT-barycenter consensus outperforms an independent learners baseline in convergence speed and final coordination success.

Keywords

Cite

@article{arxiv.2506.12497,
  title  = {Wasserstein-Barycenter Consensus for Cooperative Multi-Agent Reinforcement Learning},
  author = {Ali Baheri},
  journal= {arXiv preprint arXiv:2506.12497},
  year   = {2025}
}
R2 v1 2026-07-01T03:17:45.089Z