中文
相关论文

相关论文: Global Convergence of Policy Gradient for Sequenti…

200 篇论文

In Stackelberg v/s Stackelberg games a collection of leaders compete in a Nash game constrained by the equilibrium conditions of another Nash game amongst the followers. The resulting equilibrium problems are plagued by the nonuniqueness of…

最优化与控制 · 数学 2016-11-18 Ankur A. Kulkarni , Uday V. Shanbhag

We consider zero-sum stochastic games with finite state and action spaces, perfect information, mean payoff criteria, without any irreducibility assumption on the Markov chains associated to strategies (multichain games). The value of such…

最优化与控制 · 数学 2012-08-03 Marianne Akian , Jean Cochet-Terrasson , Sylvie Detournay , Stéphane Gaubert

In this work, we propose, for the first time, a reinforcement learning framework specifically designed for zero-sum linear-quadratic stochastic differential games. This approach offers a generalized solution for scenarios in which accurate…

最优化与控制 · 数学 2026-02-10 Yiyuan Wang

This paper addresses a Stackelberg stochastic linear-quadratic (LQ) differential game under closed-loop information, a problem inherently time-inconsistent. Existing approaches rely on solving two coupled Hamilton-Jacobi-Bellman (HJB)…

最优化与控制 · 数学 2026-04-27 Qi Lü , Bowen Ma , Hanxiao Wang

In this paper, we study finite-agent linear-quadratic games on graphs. Specifically, we propose a comprehensive framework that extends the existing literature by incorporating heterogeneous and interpretable player interactions. Compared to…

最优化与控制 · 数学 2025-11-19 Ruimeng Hu , Jihao Long , Haosheng Zhou

We study generalized games with full row rank equality constraints and we provide a strikingly simple proof of strong monotonicity of the associated KKT operator. This allows us to show linear convergence to a variational equilibrium of the…

最优化与控制 · 数学 2023-04-20 Mattia Bianchi , Emilio Benenati , Sergio Grammatico

We introduce a contractive abstract dynamic programming framework and related policy iteration algorithms, specifically designed for sequential zero-sum games and minimax problems with a general structure. Aside from greater generality, the…

计算机科学与博弈论 · 计算机科学 2021-10-22 Dimitri Bertsekas

In this paper, the known deterministic linear-quadratic Stackelberg game is revisited, whose open-loop Stackelberg solution actually possesses the nature of time inconsistency. To handle this time inconsistency, {a two-tier game framework…

最优化与控制 · 数学 2022-03-09 Yuan-Hua Ni , Liping Liu , Xinzhen Zhang

Policy gradient methods enjoy strong practical performance in numerous tasks in reinforcement learning. Their theoretical understanding in multiagent settings, however, remains limited, especially beyond two-player competitive and potential…

计算机科学与博弈论 · 计算机科学 2023-12-22 Ioannis Anagnostides , Ioannis Panageas , Gabriele Farina , Tuomas Sandholm

$ $This paper addresses the inverse problem for Linear-Quadratic (LQ) nonzero-sum $N$-player differential games, where the goal is to learn parameters of an unknown cost function for the game, called observed, given the demonstrated…

最优化与控制 · 数学 2024-10-28 Emin Martirosyan , Ming Cao

As quantum processors advance, the emergence of large-scale decentralized systems involving interacting quantum-enabled agents is on the horizon. Recent research efforts have explored quantum versions of Nash and correlated equilibria as…

计算机科学与博弈论 · 计算机科学 2024-12-18 Wayne Lin , Georgios Piliouras , Ryann Sim , Antonios Varvitsiotis

Game-theoretic models of learning are a powerful set of models that optimize multi-objective architectures. Among these models are zero-sum architectures that have inspired adversarial learning frameworks. An important shortcoming of these…

机器学习 · 计算机科学 2020-06-09 Ari Azarafrooz

A growing line of work reframes preference-based fine-tuning of large language models game-theoretically: Nash Learning from Human Feedback (NLHF) recasts the problem as a zero-sum game over policies. However, optimization is over expected…

计算机科学与博弈论 · 计算机科学 2026-05-14 Max Horwitz , Jake Gonzales , Eric Mazumdar , Lillian J. Ratliff

In this paper, Nash equilibrium seeking among a network of players is considered. Different from many existing works on Nash equilibrium seeking in non-cooperative games, the players considered in this paper cannot directly observe the…

最优化与控制 · 数学 2017-03-28 Maojiao Ye , Guoqiang Hu

This paper investigates a two-person non-homogeneous linear-quadratic stochastic differential game (LQ-SDG, for short) in an infinite horizon for a system regulated by a time-invariant Markov chain. Both non-zero-sum and zero-sum LQ-SDG…

最优化与控制 · 数学 2024-08-26 Fan Wu , Xun Li , Jie Xiong , Xin Zhang

Consider a two-player zero-sum stochastic game where the transition function can be embedded in a given feature space. We propose a two-player Q-learning algorithm for approximating the Nash equilibrium strategy via sampling. The algorithm…

机器学习 · 计算机科学 2019-06-04 Zeyu Jia , Lin F. Yang , Mengdi Wang

This paper presents a model-free approximation for the Hessian of the performance of deterministic policies to use in the context of Reinforcement Learning based on Quasi-Newton steps in the policy parameters. We show that the approximate…

机器学习 · 计算机科学 2022-03-29 Arash Bahari Kordabad , Hossein Nejatbakhsh Esfahani , Wenqi Cai , Sebastien Gros

Multi-robot coordination often exhibits hierarchical structure, with some robots' decisions depending on the planned behaviors of others. While game theory provides a principled framework for such interactions, existing solvers struggle to…

计算机科学与博弈论 · 计算机科学 2026-05-18 Hamzah Khan , Dong Ho Lee , Jingqi Li , Tianyu Qiu , Christian Ellis , Jesse Milzman , Wesley Suttle , David Fridovich-Keil

We provide a distributed algorithm to learn a Nash equilibrium in a class of non-cooperative games with strongly monotone mappings and unconstrained action sets. Each player has access to her own smooth local cost function and can…

最优化与控制 · 数学 2019-07-17 Tatiana Tatarenko , Angelia Nedich

In this paper, we first address a linear quadratic mean-field game problem with a leader-follower structure. By adopting a Riccati-type approach, we show how one can obtain a state-feedback representation of the pairs of strategies which…

系统与控制 · 电气工程与系统科学 2023-02-21 Samir Aberkane , Vasile Dragan