中文
相关论文

相关论文: A Generalized Minimax Q-learning Algorithm for Two…

200 篇论文

We consider the cyber-physical security of parallel server systems, which is relevant for a variety of engineering applications such as networking, manufacturing, and transportation. These systems rely on feedback control and may thus be…

系统与控制 · 电气工程与系统科学 2025-07-18 Yuzhen Zhan , Li Jin

We develop a method based on computer algebra systems to represent the mutual pure strategy best-response dynamics of symmetric two-player, two-action repeated games played by players with a one-period memory. We apply this method to the…

动力系统 · 数学 2022-10-04 Janusz M Meylahn , Lars Janssen

In this paper, we formulate a two-player zero-sum game under dynamic constraints defined by hybrid dynamical equations. The game consists of a min-max problem involving a cost functional that depends on the actions and resulting solutions…

最优化与控制 · 数学 2025-05-20 Santiago J. Leudo , Ricardo G. Sanfelice

We revisit the problem of learning in two-player zero-sum Markov games, focusing on developing an algorithm that is uncoupled, convergent, and rational, with non-asymptotic convergence rates. We start from the case of stateless matrix game…

计算机科学与博弈论 · 计算机科学 2023-11-10 Yang Cai , Haipeng Luo , Chen-Yu Wei , Weiqiang Zheng

We study a two-player, zero-sum, stochastic game with incomplete information on one side in which the players are allowed to play more and more frequently. The informed player observes the realization of a Markov chain on which the payoffs…

最优化与控制 · 数学 2013-07-15 Pierre Cardaliaguet , Catherine Rainer , Dinah Rosenberg , Nicolas Vieille

Two-player complete-information game trees are perhaps the simplest possible setting for studying general-sum games and the computational problem of finding equilibria. These games admit a simple bottom-up algorithm for finding subgame…

计算机科学与博弈论 · 计算机科学 2012-07-02 Michael L. Littman , Nishkam Ravi , Arjun Talwar , Martin Zinkevich

This paper studies policy optimization algorithms for multi-agent reinforcement learning. We begin by proposing an algorithm framework for two-player zero-sum Markov Games in the full-information setting, where each iteration consists of a…

机器学习 · 计算机科学 2022-07-26 Runyu Zhang , Qinghua Liu , Huan Wang , Caiming Xiong , Na Li , Yu Bai

We investigate zero-sum turn-based two-player stochastic games in which the objective of one player is to maximize the amount of rewards obtained during a play, while the other aims at minimizing it. We focus on games in which the minimizer…

计算机科学中的逻辑 · 计算机科学 2022-05-20 Pablo F. Castro , Pedro R. D'Argenio , Luciano Putruele , Ramiro Demasi

2-TBSG is a two-player game model which aims to find Nash equilibriums and is widely utilized in reinforced learning and AI. Inspired by the fact that the simplex method for solving the deterministic discounted Markov decision processes…

计算机科学与博弈论 · 计算机科学 2019-06-11 Zeyu Jia , Zaiwen Wen , Yinyu Ye

We examine the problem of the existence of optimal deterministic stationary strategiesintwo-players antagonistic (zero-sum) perfect information stochastic games with finitely many states and actions.We show that the existenceof such…

计算机科学与博弈论 · 计算机科学 2016-11-28 Hugo Gimbert , Wieslaw Zielonka

In this paper, we propose a provably convergent and practical framework for multi-objective reinforcement learning with max-min criterion. From a game-theoretic perspective, we reformulate max-min multi-objective reinforcement learning as a…

机器学习 · 计算机科学 2025-10-24 Woohyeon Byeon , Giseung Park , Jongseong Chae , Amir Leshem , Youngchul Sung

We introduce two-level discounted games played by two players on a perfect-information stochastic game graph. The upper level game is a discounted game and the lower level game is an undiscounted reachability game. Two-level games model…

计算机科学中的逻辑 · 计算机科学 2010-06-09 Krishnendu Chatterjee , Rupak Majumdar

A new solution concept for two-player zero-sum matrix games with multi-dimensional payoff is introduced. It is based on extensions of vector orders in K-dimensional spaces to order relations in their power sets, so-called set relations, and…

最优化与控制 · 数学 2017-01-31 Andreas H. Hamel , Andreas Loehne

We design and analyze minimax-optimal algorithms for online linear optimization games where the player's choice is unconstrained. The player strives to minimize regret, the difference between his loss and the loss of a post-hoc benchmark…

机器学习 · 计算机科学 2013-02-12 H. Brendan McMahan

Strategy iteration is a technique frequently used for two-player games in order to determine the winner or compute payoffs, but to the best of our knowledge no general framework for strategy iteration has been considered. Inspired by…

计算机科学中的逻辑 · 计算机科学 2022-12-14 Paolo Baldan , Richard Eggert , Barbara König , Tommaso Padoan

Semi-Markov model is one of the most general models for stochastic dynamic systems. This paper deals with a two-person zero-sum game for semi-Markov processes. We focus on the expected discounted payoff criterion with state-action-dependent…

计算机科学与博弈论 · 计算机科学 2021-03-09 Zhihui Yu , Xianping Guo , Li Xia

The problem of two-player zero-sum Markov games has recently attracted increasing interests in theoretical studies of multi-agent reinforcement learning (RL). In particular, for finite-horizon episodic Markov decision processes (MDPs), it…

机器学习 · 计算机科学 2024-06-07 Songtao Feng , Ming Yin , Yu-Xiang Wang , Jing Yang , Yingbin Liang

The focus of this paper is a Bayesian framework for solving a class of problems termed multi-agent inverse reinforcement learning (MIRL). Compared to the well-known inverse reinforcement learning (IRL) problem, MIRL is formalized in the…

计算机科学与博弈论 · 计算机科学 2019-07-31 Xiaomin Lin , Peter A. Beling , Randy Cogill

We present a novel variant of fictitious play dynamics combining classical fictitious play with Q-learning for stochastic games and analyze its convergence properties in two-player zero-sum stochastic games. Our dynamics involves players…

计算机科学与博弈论 · 计算机科学 2022-06-03 Muhammed O. Sayin , Francesca Parise , Asuman Ozdaglar

In this article we analyze a partial-information Nash Q-learning algorithm for a general 2-player stochastic game. Partial information refers to the setting where a player does not know the strategy or the actions taken by the opposing…

计算机科学与博弈论 · 计算机科学 2023-02-22 Negash Medhin , Andrew Papanicolaou , Marwen Zrida