中文
相关论文

相关论文: Global Convergence of Policy Gradient for Sequenti…

200 篇论文

We present a midpoint policy iteration algorithm to solve linear quadratic optimal control problems in both model-based and model-free settings. The algorithm is a variation of Newton's method, and we show that in the model-based setting it…

最优化与控制 · 数学 2022-02-16 Benjamin Gravell , Iman Shames , Tyler Summers

Nash equilibrium has long been a desired solution concept in multi-player games, especially for those on continuous strategy spaces, which have attracted a rapidly growing amount of interests due to advances in research applications such as…

计算机科学与博弈论 · 计算机科学 2019-10-29 Zehao Dou , Xiang Yan , Dongge Wang , Xiaotie Deng

We consider seeking generalized Nash equilibria (GNE) for noncooperative games with coupled nonlinear constraints over networks. We first revisit a well-known gradientplay dynamics from a passivity-based perspective, and address that the…

最优化与控制 · 数学 2024-08-23 Weijian Li , Lacra Pavel

The concept of leader--follower (or Stackelberg) equilibrium plays a central role in a number of real--world applications of game theory. While the case with a single follower has been thoroughly investigated, results with multiple…

计算机科学与博弈论 · 计算机科学 2017-07-10 Nicola Basilico , Stefano Coniglio , Nicola Gatti

This paper proposes a new method for finding closed-loop saddle points in zero-sum linear-quadratic stochastic differential games by decoupling their inherent structure. Specifically, we develop a nested iterative scheme that constructs a…

最优化与控制 · 数学 2025-12-10 Yiyuan Wang

We consider a distributed stochastic approximation (SA) scheme for computing an equilibrium of a stochastic Nash game. Standard SA schemes employ diminishing steplength sequences that are square summable but not summable. Such requirements…

最优化与控制 · 数学 2013-03-20 Farzad Yousefian , Angelia Nedich , Uday V. Shanbhag

This paper aims to formulate and study the inverse problem of non-cooperative linear quadratic games: Given a profile of control strategies, find cost parameters for which this profile of control strategies is Nash. We formulate the problem…

最优化与控制 · 数学 2022-07-14 Yunhan Huang , Tao Zhang , Quanyan Zhu

In game theory, the concept of Nash equilibrium reflects the collective stability of some individual strategies chosen by selfish agents. The concept pertains to different classes of games, e.g. the sequential games, where the agents play…

逻辑 · 数学 2015-07-01 Stephane Le Roux

Contemporary applications of machine learning in two-team e-sports and the superior expressivity of multi-agent generative adversarial networks raise important and overlooked theoretical questions regarding optimization in two-team games.…

计算机科学与博弈论 · 计算机科学 2023-04-18 Fivos Kalogiannis , Ioannis Panageas , Emmanouil-Vasileios Vlatakis-Gkaragkounis

We consider a class of non-cooperative N-player non-zero-sum stochastic differential games with singular controls, in which each player can affect a linear stochastic differential equation in order to minimize a cost functional which is…

最优化与控制 · 数学 2023-04-19 Jodi Dianetti

This paper focuses on a kind of linear quadratic non-zero sum differential game driven by backward stochastic differential equation with asymmetric information, which is a natural continuation of Wang and Yu [IEEE TAC (2010) 55: 1742-1747,…

最优化与控制 · 数学 2017-03-06 Guangchen Wang , Hua Xiao , Jie Xiong

We introduce a novel class of Nash equilibrium seeking dynamics for non-cooperative games with a finite number of players, where the convergence to the Nash equilibrium is bounded by a KL function with a settling time that can be upper…

最优化与控制 · 数学 2020-12-25 Jorge I. Poveda , Miroslav Krstic , Tamer Basar

We introduce a stochastic learning process called the dampened gradient approximation process. While learning models have almost exclusively focused on finite games, in this paper we design a learning process for games with continuous…

计算机科学与博弈论 · 计算机科学 2018-07-02 Sebastian Bervoets , Mario Bravo , Mathieu Faure

In this paper, we present exploitability descent, a new algorithm to compute approximate equilibria in two-player zero-sum extensive-form games with imperfect information, by direct policy optimization against worst-case opponents. We prove…

This paper investigates a robust incentive Stackelberg stochastic differential game problem for a linear-quadratic mean field system, where the model uncertainty appears in the drift term of the leader's state equation. Moreover, both the…

最优化与控制 · 数学 2026-03-31 Na Xiang , Jingtao Shi

We consider the problem of efficiently learning to play single-leader multi-follower Stackelberg games when the leader lacks knowledge of the lower-level game. Such games arise in hierarchical decision-making problems involving…

系统与控制 · 电气工程与系统科学 2025-12-11 Anna Maddux , Marko Maljkovic , Nikolas Geroliminis , Maryam Kamgarpour

Policy gradient algorithms have been widely applied to Markov decision processes and reinforcement learning problems in recent years. Regularization with various entropy functions is often used to encourage exploration and improve…

机器学习 · 计算机科学 2023-06-09 Haoya Li , Samarth Gupta , Hsiangfu Yu , Lexing Ying , Inderjit Dhillon

We prove that differential Nash equilibria are generic amongst local Nash equilibria in continuous zero-sum games. That is, there exists an open-dense subset of zero-sum games for which local Nash equilibria are non-degenerate differential…

计算机科学与博弈论 · 计算机科学 2020-02-05 Eric Mazumdar , Lillian Ratliff

We consider infinite-horizon discounted Markov decision processes and study the convergence rates of the natural policy gradient (NPG) and the Q-NPG methods with the log-linear policy class. Using the compatible function approximation…

机器学习 · 计算机科学 2023-02-22 Rui Yuan , Simon S. Du , Robert M. Gower , Alessandro Lazaric , Lin Xiao

We extend the formalism of Conjectural Variations games to Stackelberg games involving multiple leaders and a single follower. To solve these nonconvex games, a common assumption is that the leaders compute their strategies having perfect…

计算机科学与博弈论 · 计算机科学 2025-07-24 Francesco Morri , Hélène Le Cadre , Luce Brotcorne
‹ 上一页 1 8 9 10 下一页 ›