中文
相关论文

相关论文: Learning Zero-Sum Linear Quadratic Games with Impr…

200 篇论文

There has been significant recent progress in algorithms for approximation of Nash equilibrium in large two-player zero-sum imperfect-information games and exact computation of Nash equilibrium in multiplayer strategic-form games. While…

计算机科学与博弈论 · 计算机科学 2025-10-01 Sam Ganzfried

We consider the problem of optimal tracking control of unknown discrete-time nonlinear nonzero-sum games. The related state-of-art literature is mostly focused on Policy Iteration algorithms and multiple neural network approximation, which…

系统与控制 · 电气工程与系统科学 2023-11-07 Alexandros Tanzanakis , John Lygeros

Motivated by Generative Adversarial Networks, we study the computation of Nash equilibrium in concave network zero-sum games (NZSGs), a multiplayer generalization of two-player zero-sum games first proposed with linear payoffs. Extending…

机器学习 · 计算机科学 2020-07-13 Amit Kadan , Hu Fu

We develop a flexible stochastic approximation framework for analyzing the long-run behavior of learning in games (both continuous and finite). The proposed analysis template incorporates a wide array of popular learning algorithms,…

计算机科学与博弈论 · 计算机科学 2023-07-04 Panayotis Mertikopoulos , Ya-Ping Hsieh , Volkan Cevher

We develop a novel primal-dual algorithm to solve a class of nonsmooth and nonlinear compositional convex minimization problems, which covers many existing and brand-new models as special cases. Our approach relies on a combination of a new…

最优化与控制 · 数学 2021-04-20 Yuzixuan Zhu , Deyi Liu , Quoc Tran-Dinh

We study multi-agent general-sum Markov games with nonlinear function approximation. We focus on low-rank Markov games whose transition matrix admits a hidden low-rank structure on top of an unknown non-linear representation. The goal is to…

机器学习 · 计算机科学 2022-11-01 Chengzhuo Ni , Yuda Song , Xuezhou Zhang , Chi Jin , Mengdi Wang

Safety is critical during human-robot interaction. But -- because people are inherently unpredictable -- it is often difficult for robots to plan safe behaviors. Instead of relying on our ability to anticipate humans, here we identify robot…

机器人学 · 计算机科学 2026-04-07 Benjamin A. Christie , Dylan P. Losey

We address black-box convex optimization problems, where the objective and constraint functions are not explicitly known but can be sampled within the feasible set. The challenge is thus to generate a sequence of feasible points converging…

最优化与控制 · 数学 2022-11-08 Baiwei Guo , Yuning Jiang , Maryam Kamgarpour , Giancarlo Ferrari-Trecate

We investigate a class of zero-sum linear-quadratic stochastic differential games on a finite time horizon governed by multiscale state equations. The multiscale nature of the problem can be leveraged to reformulate the associated…

最优化与控制 · 数学 2020-11-19 Beniamin Goldys , James Yang , Zhou Zhou

This paper studies accelerated algorithms for Q-learning. We propose an acceleration scheme by incorporating the historical iterates of the Q-function. The idea is conceptually inspired by the momentum-based acceleration methods in the…

系统与控制 · 电气工程与系统科学 2019-10-28 Bowen Weng , Lin Zhao , Huaqing Xiong , Wei Zhang

This paper formulates and studies a linear quadratic (LQ for short) game problem governed by linear stochastic Volterra integral equation. Sufficient and necessary condition of the existence of saddle points for this problem are derived. As…

概率论 · 数学 2010-05-31 Tianxiao Wang , Yufeng Shi

Inspired by REINFORCE, we introduce a novel receding-horizon algorithm for the Linear Quadratic Regulator (LQR) problem with unknown dynamics. Unlike prior methods, our algorithm avoids reliance on two-point gradient estimates while…

最优化与控制 · 数学 2025-10-07 Amirreza Neshaei Moghaddam , Alex Olshevsky , Bahman Gharesifard

As quantum processors advance, the emergence of large-scale decentralized systems involving interacting quantum-enabled agents is on the horizon. Recent research efforts have explored quantum versions of Nash and correlated equilibria as…

计算机科学与博弈论 · 计算机科学 2024-12-18 Wayne Lin , Georgios Piliouras , Ryann Sim , Antonios Varvitsiotis

We consider constrained linear-quadratic dynamic games arising in autonomous vehicle platooning, intersection crossing and other cooperative driving scenarios. Infinite-horizon Nash equilibria are reformulated as receding-horizon affine…

最优化与控制 · 数学 2026-05-05 Reza Rahimi Baghbadorani , Sergio Grammatico

With the outstanding performance of policy gradient (PG) method in the reinforcement learning field, the convergence theory of it has aroused more and more interest recently. Meanwhile, the significant importance and abundant theoretical…

最优化与控制 · 数学 2024-04-19 Xinpei Zhang , Guangyan Jia

We study the problem of learning Nash equilibria in offline two-player zero-sum Markov games. While existing approaches often rely on explicit pessimism to address distribution shift, we show that KL regularization alone suffices to…

机器学习 · 计算机科学 2026-05-14 Claire Chen , Yuheng Zhang , Xinyu Liu , Zixuan Xie , Shuze Daniel Liu , Nan Jiang

Dynamic games can be an effective approach to modeling interactive behavior between multiple non-cooperative agents and they provide a theoretical framework for simultaneous prediction and control in such scenarios. In this work, we propose…

系统与控制 · 电气工程与系统科学 2022-09-19 Edward L. Zhu , Francesco Borrelli

We consider two classes of constrained finite state-action stochastic games. First, we consider a two player nonzero sum single controller constrained stochastic game with both average and discounted cost criterion. We consider the same…

最优化与控制 · 数学 2012-06-11 Vikas Vikram Singh , N. Hemachandra

Last-iterate convergence of learning dynamics in games has attracted significant recent attention. In two-player zero-sum games with bandit feedback, where only the loss of the selected action pair is observed, Fiegel et al. (2025) show a…

机器学习 · 计算机科学 2026-05-12 Soumita Hait , Ping Li , Haipeng Luo , Mengxiao Zhang

We consider online learning in multi-player smooth monotone games. Existing algorithms have limitations such as (1) being only applicable to strongly monotone games; (2) lacking the no-regret guarantee; (3) having only asymptotic or slow…

机器学习 · 计算机科学 2023-09-06 Yang Cai , Weiqiang Zheng