English

Bandit Learning in Convex Non-Strictly Monotone Games

Optimization and Control 2023-08-17 v5

Abstract

We address learning Nash equilibria in convex games under the payoff information setting. We consider the case in which the game pseudo-gradient is monotone but not necessarily strictly monotone. This relaxation of strict monotonicity enables application of learning algorithms to a larger class of games, such as, for example, a zero-sum game with a merely convex-concave cost function. We derive an algorithm whose iterates provably converge to the least-norm Nash equilibrium in this setting. {From the perspective of a single player using the proposed algorithm, we view the game as an instance of online optimization}. Through this lens, we quantify the regret rate of the algorithm and provide an approach to choose the algorithm's parameters to minimize the regret rate.

Keywords

Cite

@article{arxiv.2009.04258,
  title  = {Bandit Learning in Convex Non-Strictly Monotone Games},
  author = {Tatiana Tatarenko and Maryam Kamgarpour},
  journal= {arXiv preprint arXiv:2009.04258},
  year   = {2023}
}

Comments

arXiv admin note: text overlap with arXiv:1904.01882

R2 v1 2026-06-23T18:24:54.176Z