English

Learning to Play General-Sum Games Against Multiple Boundedly Rational Agents

Machine Learning 2022-12-21 v3

Abstract

We study the problem of training a principal in a multi-agent general-sum game using reinforcement learning (RL). Learning a robust principal policy requires anticipating the worst possible strategic responses of other agents, which is generally NP-hard. However, we show that no-regret dynamics can identify these worst-case responses in poly-time in smooth games. We propose a framework that uses this policy evaluation method for efficiently learning a robust principal policy using RL. This framework can be extended to provide robustness to boundedly rational agents too. Our motivating application is automated mechanism design: we empirically demonstrate our framework learns robust mechanisms in both matrix games and complex spatiotemporal games. In particular, we learn a dynamic tax policy that improves the welfare of a simulated trade-and-barter economy by 15%, even when facing previously unseen boundedly rational RL taxpayers.

Keywords

Cite

@article{arxiv.2106.05492,
  title  = {Learning to Play General-Sum Games Against Multiple Boundedly Rational Agents},
  author = {Eric Zhao and Alexander R. Trott and Caiming Xiong and Stephan Zheng},
  journal= {arXiv preprint arXiv:2106.05492},
  year   = {2022}
}

Comments

15 pages, 6 figures. Appearing at the Thirty-seventh AAAI Conference on Artificial Intelligence (AAAI 2023)

R2 v1 2026-06-24T03:02:25.950Z