English
Related papers

Related papers: Global Convergence of Policy Gradient for Sequenti…

200 papers

We examine global non-asymptotic convergence properties of policy gradient methods for multi-agent reinforcement learning (RL) problems in Markov potential games (MPG). To learn a Nash equilibrium of an MPG in which the size of state space…

Machine Learning · Computer Science 2022-08-08 Dongsheng Ding , Chen-Yu Wei , Kaiqing Zhang , Mihailo R. Jovanović

A growing body of work in game theory extends the traditional Stackelberg game to settings with one leader and multiple followers who play a Nash equilibrium. Standard approaches for computing equilibria in these games reformulate the…

Computer Science and Game Theory · Computer Science 2021-12-07 Kai Wang , Lily Xu , Andrew Perrault , Michael K. Reiter , Milind Tambe

We explore reinforcement learning methods for finding the optimal policy in the linear quadratic regulator (LQR) problem. In particular, we consider the convergence of policy gradient methods in the setting of known and unknown parameters.…

Machine Learning · Computer Science 2021-06-25 Ben Hambly , Renyuan Xu , Huining Yang

We study infinite-horizon discounted two-player zero-sum Markov games, and develop a decentralized algorithm that provably converges to the set of Nash equilibria under self-play. Our algorithm is based on running an Optimistic Gradient…

Machine Learning · Computer Science 2021-07-08 Chen-Yu Wei , Chung-Wei Lee , Mengxiao Zhang , Haipeng Luo

There has been substantial recent progress on the theoretical understanding of model-free approaches to Linear Quadratic Regulator (LQR) problems. Much attention has been devoted to the special case when the goal is to drive the state close…

Optimization and Control · Mathematics 2021-04-13 Zhaolin Ren , Aoxiao Zhong , Na Li

We analyze best response dynamics for finding a Nash equilibrium of an infinite horizon zero-sum stochastic linear quadratic dynamic game (LQDG) with partial and asymmetric information. We derive explicit expressions for each player's best…

Systems and Control · Electrical Eng. & Systems 2025-09-03 Yuxiang Guan , Iman Shames , Tyler H. Summers

We consider the continuous-time Linear-Quadratic-Regulator (LQR) problem in terms of optimizing a real-valued matrix function over the set of feedback gains. The results developed are in parallel to those in Bu et al. [1] for discrete-time…

Systems and Control · Electrical Eng. & Systems 2020-06-17 Jingjing Bu , Afshin Mesbahi , Mehran Mesbahi

In this paper, we investigate a class of nonzero-sum dynamic stochastic games, where players have linear dynamics and quadratic cost functions. The players are coupled in both dynamics and cost through a linear regression (weighted average)…

Optimization and Control · Mathematics 2020-10-20 Jalal Arabneydi , Amir G. Aghdam , Roland P. Malhamé

Policy gradient algorithms are widely used in reinforcement learning and belong to the class of approximate dynamic programming methods. This paper studies two key policy gradient algorithms, the Natural Policy Gradient and the Gauss-Newton…

Systems and Control · Electrical Eng. & Systems 2026-05-11 Bowen Song , Sebastien Gros , Andrea Iannelli

We study the problem of computing an approximate Nash equilibrium of continuous-action game without access to gradients. Such game access is common in reinforcement learning settings, where the environment is typically treated as a black…

Computer Science and Game Theory · Computer Science 2023-08-30 Carlos Martin , Tuomas Sandholm

We derive sublinear-time quantum algorithms for computing the Nash equilibrium of two-player zero-sum games, based on efficient Gibbs sampling methods. We are able to achieve speed-ups for both dense and sparse payoff matrices at the cost…

Quantum Physics · Physics 2019-04-08 Joran van Apeldoorn , András Gilyén

Multi-agent games in dynamic nonlinear settings are challenging due to the time-varying interactions among the agents and the non-stationarity of the (potential) Nash equilibria. In this paper we consider model-free games, where agent…

Systems and Control · Electrical Eng. & Systems 2025-09-24 Eduardo Sebastián , Maitrayee Keskar , Eeman Iqbal , Eduardo Montijano , Carlos Sagüés , Nikolay Atanasov

We consider a class of smooth $N$-player noncooperative games, where player objectives are expectation-valued and potentially nonconvex. In such a setting, we consider the largely open question of efficiently computing a suitably defined…

Optimization and Control · Mathematics 2025-05-23 Zhuoyu Xiao , Uday V. Shanbhag

We study the convergence to local Nash equilibria of gradient methods for two-player zero-sum differentiable games. It is well-known that such dynamics converge locally when $S \succ 0$ and may diverge when $S=0$, where $S\succeq 0$ is the…

Optimization and Control · Mathematics 2023-11-08 Guillaume Wang , Lénaïc Chizat

We contribute the first provable guarantees of global convergence to Nash equilibria (NE) in two-player zero-sum convex Markov games (cMGs) by using independent policy gradient methods. Convex Markov games, recently defined by Gemp et al.…

Computer Science and Game Theory · Computer Science 2025-06-23 Fivos Kalogiannis , Emmanouil-Vasileios Vlatakis-Gkaragkounis , Ian Gemp , Georgios Piliouras

This paper considers a distributed Nash equilibrium seeking problem, where the players only have partial access to other players' actions, such as their neighbors' actions. Thus, the players are supposed to communicate with each other to…

Optimization and Control · Mathematics 2020-03-31 Yipeng Pang , Guoqiang Hu

We study Stackelberg equilibria in finitely repeated games, where the leader commits to a strategy that picks actions in each round and can be adaptive to the history of play (i.e. they commit to an algorithm). In particular, we study…

Computer Science and Game Theory · Computer Science 2024-03-08 Natalie Collina , Eshwar Ram Arunachaleswaran , Michael Kearns

This paper is concerned with a linear-quadratic (LQ) Stackelberg mean field games of backward-forward stochastic systems, involving a backward leader and a substantial number of forward followers. The leader initiates by providing its…

Optimization and Control · Mathematics 2024-06-28 Wenyu Cong , Jingtao Shi

Direct policy gradient methods for reinforcement learning and continuous control problems are a popular approach for a variety of reasons: 1) they are easy to implement without explicit knowledge of the underlying model 2) they are an…

Machine Learning · Computer Science 2019-03-26 Maryam Fazel , Rong Ge , Sham M. Kakade , Mehran Mesbahi

Motivated by Generative Adversarial Networks, we study the computation of Nash equilibrium in concave network zero-sum games (NZSGs), a multiplayer generalization of two-player zero-sum games first proposed with linear payoffs. Extending…

Machine Learning · Computer Science 2020-07-13 Amit Kadan , Hu Fu