中文
相关论文

相关论文: Towards General Function Approximation in Zero-Sum…

200 篇论文

The objective of this paper is to investigate the finite-time analysis of a Q-learning algorithm applied to two-player zero-sum Markov games. Specifically, we establish a finite-time analysis of both the minimax Q-learning algorithm and the…

系统与控制 · 电气工程与系统科学 2023-06-13 Donghwan Lee

This paper is concerned with two-person dynamic zero-sum games. Let games for some family have common dynamics, running costs and capabilities of players, and let these games differ in densities only. We show that the Dynamic Programming…

最优化与控制 · 数学 2017-09-26 Dmitry Khlopin

In this paper, we consider the problem of optimization and learning for constrained and multi-objective Markov decision processes, for both discounted rewards and expected average rewards. We formulate the problems as zero-sum games where…

最优化与控制 · 数学 2021-03-05 Ather Gattami , Qinbo Bai , Vaneet Agarwal

We study a class of zero-sum stochastic games between a stopper and a singular-controller, previously considered in [Bovo and De Angelis (2025)]. The underlying singularly-controlled dynamics takes values in…

最优化与控制 · 数学 2025-06-25 Andrea Bovo , Alessandro Milazzo

We study the problem of repeated play in a zero-sum game in which the payoff matrix may change, in a possibly adversarial fashion, on each round; we call these Online Matrix Games. Finding the Nash Equilibrium (NE) of a two player zero-sum…

机器学习 · 计算机科学 2020-04-06 Adrian Rivera Cardoso , Jacob Abernethy , He Wang , Huan Xu

We contribute the first provable guarantees of global convergence to Nash equilibria (NE) in two-player zero-sum convex Markov games (cMGs) by using independent policy gradient methods. Convex Markov games, recently defined by Gemp et al.…

计算机科学与博弈论 · 计算机科学 2025-06-23 Fivos Kalogiannis , Emmanouil-Vasileios Vlatakis-Gkaragkounis , Ian Gemp , Georgios Piliouras

Zero-sum Dynkin games under Poisson constraints, where players can only stop at the event times of a Poisson process, have been studied widely in the recent literature. The constraint can be modelled in two ways: either both players share…

最优化与控制 · 数学 2025-12-09 David Hobson , Gechun Liang , Edward Wang

Two-player zero-sum games are a well-established model for synthesising controllers that optimise some performance criterion. In such games one player represents the controller, while the other describes the (adversarial) environment, and…

计算机科学与博弈论 · 计算机科学 2010-06-04 Marta Kwiatkowska , Gethin Norman , Ashutosh Trivedi

Stochastic games generalize Markov decision processes (MDPs) to a multiagent setting by allowing the state transitions to depend jointly on all player actions, and having rewards determined by multiplayer matrix games at each state. We…

计算机科学与博弈论 · 计算机科学 2013-01-18 Michael Kearns , Yishay Mansour , Satinder Singh

Computing approximate Nash equilibria in multi-player general-sum Markov games is a computationally intractable task. However, multi-player Markov games with certain cooperative or competitive structures might circumvent this…

计算机科学与博弈论 · 计算机科学 2023-08-17 Zailin Ma , Jiansheng Yang , Zhihua Zhang

This paper considers a new class of deterministic finite-time horizon, two-player, zero-sum differential games (DGs) in which the maximizing player is allowed to take continuous and impulse controls whereas the minimizing player is allowed…

最优化与控制 · 数学 2022-12-21 Brahim El Asri , Hafid Lalioui

We examine global non-asymptotic convergence properties of policy gradient methods for multi-agent reinforcement learning (RL) problems in Markov potential games (MPG). To learn a Nash equilibrium of an MPG in which the size of state space…

机器学习 · 计算机科学 2022-08-08 Dongsheng Ding , Chen-Yu Wei , Kaiqing Zhang , Mihailo R. Jovanović

The works of (Daskalakis et al., 2009, 2022; Jin et al., 2022; Deng et al., 2023) indicate that computing Nash equilibria in multi-player Markov games is a computationally hard task. This fact raises the question of whether or not…

计算机科学与博弈论 · 计算机科学 2023-05-30 Fivos Kalogiannis , Ioannis Panageas

A basic question for zero-sum repeated games consists in determining whether the mean payoff per time unit is independent of the initial state. In the special case of "zero-player" games, i.e., of Markov chains equipped with additive…

最优化与控制 · 数学 2015-10-20 Marianne Akian , Stéphane Gaubert , Antoine Hochart

This paper develops an algorithm for upper- and lower-bounding the value function for a class of linear time-varying games subject to convex control sets. In particular, a two-player zero-sum differential game is considered where the…

最优化与控制 · 数学 2025-03-12 Vincent Liu , Chris Manzie , Peter M. Dower

Finding equilibria via gradient play in competitive multi-agent games has been attracting a growing amount of attention in recent years, with emphasis on designing efficient strategies where the agents operate in a decentralized and…

计算机科学与博弈论 · 计算机科学 2022-11-17 Ruicheng Ao , Shicong Cen , Yuejie Chi

For two-person dynamic zero-sum games (both discrete and continuous settings), we investigate the limit of value functions of finite horizon games with long run average cost as the time horizon tends to infinity and the limit of value…

最优化与控制 · 数学 2017-09-26 Dmitry Khlopin

This paper studies partially observable two-person zero-sum semi-Markov games under a probability criterion, in which the system state may not be completely observed. It focuses on the probability that the accumulated rewards of player 1…

最优化与控制 · 数学 2025-08-26 Xin Wen , Li Xia , Zhihui Yu

Pursuit-evasion scenarios appear widely in robotics, security domains, and many other real-world situations. We focus on two-player pursuit-evasion games with concurrent moves, infinite horizon, and discounted rewards. We assume that the…

计算机科学与博弈论 · 计算机科学 2016-08-05 Karel Horák , Branislav Bošanský

It is common to encounter large-scale monotone inclusion problems where the objective has a finite sum structure. We develop a general framework for variance-reduced forward-backward splitting algorithms for this problem. This framework…

机器学习 · 统计学 2021-03-17 Xun Zhang , William B. Haskell , Zhisheng Ye