English

Uncoupled Learning Dynamics with $O(\log T)$ Swap Regret in Multiplayer Games

Computer Science and Game Theory 2022-10-07 v2 Machine Learning

Abstract

In this paper we establish efficient and \emph{uncoupled} learning dynamics so that, when employed by all players in a general-sum multiplayer game, the \emph{swap regret} of each player after TT repetitions of the game is bounded by O(logT)O(\log T), improving over the prior best bounds of O(log4(T))O(\log^4 (T)). At the same time, we guarantee optimal O(T)O(\sqrt{T}) swap regret in the adversarial regime as well. To obtain these results, our primary contribution is to show that when all players follow our dynamics with a \emph{time-invariant} learning rate, the \emph{second-order path lengths} of the dynamics up to time TT are bounded by O(logT)O(\log T), a fundamental property which could have further implications beyond near-optimally bounding the (swap) regret. Our proposed learning dynamics combine in a novel way \emph{optimistic} regularized learning with the use of \emph{self-concordant barriers}. Further, our analysis is remarkably simple, bypassing the cumbersome framework of higher-order smoothness recently developed by Daskalakis, Fishelson, and Golowich (NeurIPS'21).

Keywords

Cite

@article{arxiv.2204.11417,
  title  = {Uncoupled Learning Dynamics with $O(\log T)$ Swap Regret in Multiplayer Games},
  author = {Ioannis Anagnostides and Gabriele Farina and Christian Kroer and Chung-Wei Lee and Haipeng Luo and Tuomas Sandholm},
  journal= {arXiv preprint arXiv:2204.11417},
  year   = {2022}
}

Comments

To appear at NeurIPS 2022. V2 incorporates reviewers' feedback and minor corrections

R2 v1 2026-06-24T10:57:19.813Z