Uncoupled Learning Dynamics with $O(\log T)$ Swap Regret in Multiplayer Games
Abstract
In this paper we establish efficient and \emph{uncoupled} learning dynamics so that, when employed by all players in a general-sum multiplayer game, the \emph{swap regret} of each player after repetitions of the game is bounded by , improving over the prior best bounds of . At the same time, we guarantee optimal swap regret in the adversarial regime as well. To obtain these results, our primary contribution is to show that when all players follow our dynamics with a \emph{time-invariant} learning rate, the \emph{second-order path lengths} of the dynamics up to time are bounded by , a fundamental property which could have further implications beyond near-optimally bounding the (swap) regret. Our proposed learning dynamics combine in a novel way \emph{optimistic} regularized learning with the use of \emph{self-concordant barriers}. Further, our analysis is remarkably simple, bypassing the cumbersome framework of higher-order smoothness recently developed by Daskalakis, Fishelson, and Golowich (NeurIPS'21).
Cite
@article{arxiv.2204.11417,
title = {Uncoupled Learning Dynamics with $O(\log T)$ Swap Regret in Multiplayer Games},
author = {Ioannis Anagnostides and Gabriele Farina and Christian Kroer and Chung-Wei Lee and Haipeng Luo and Tuomas Sandholm},
journal= {arXiv preprint arXiv:2204.11417},
year = {2022}
}
Comments
To appear at NeurIPS 2022. V2 incorporates reviewers' feedback and minor corrections