中文
相关论文

相关论文: Policy iteration algorithm for zero-sum multichain…

200 篇论文

This paper studies policy optimization algorithms for multi-agent reinforcement learning. We begin by proposing an algorithm framework for two-player zero-sum Markov Games in the full-information setting, where each iteration consists of a…

机器学习 · 计算机科学 2022-07-26 Runyu Zhang , Qinghua Liu , Huan Wang , Caiming Xiong , Na Li , Yu Bai

We study a two-player discounted zero-sum stochastic game model for dynamic operational planning in military campaigns. At each stage, the players manage multiple commanders who order military actions on objectives that have an open line of…

计算机科学与博弈论 · 计算机科学 2024-03-04 Joseph E. McCarthy , Mathieu Dahan , Chelsea C. White

We study a finite-horizon two-person zero-sum risk-sensitive stochastic game for continuous-time Markov chains and Borel state and action spaces, in which payoff rates, transition rates and terminal reward functions are allowed to be…

最优化与控制 · 数学 2021-03-09 Junyu Zhang , Xianping Guo , Li Xia

One of the proposed solutions to the equilibrium selection problem for agents learning in repeated games is obtained via the notion of stochastic stability. Learning algorithms are perturbed so that the Markov chain underlying the learning…

计算机科学与博弈论 · 计算机科学 2012-07-09 John Wicks , Amy Greenwald

We consider the general model of zero-sum repeated games (or stochastic games with signals), and assume that one of the players is fully informed and controls the transitions of the state variable. We prove the existence of the uniform…

最优化与控制 · 数学 2009-04-20 Jérôme Renault

We consider concurrent mean-payoff games, a very well-studied class of two-player (player 1 vs player 2) zero-sum games on finite-state graphs where every transition is assigned a reward between 0 and 1, and the payoff function is the…

计算机科学与博弈论 · 计算机科学 2014-10-02 Krishnendu Chatterjee , Rasmus Ibsen-Jensen

We study a subclass of $n$-player stochastic games, namely, stochastic games with independent chains and unknown transition matrices. In this class of games, players control their own internal Markov chains whose transitions do not depend…

计算机科学与博弈论 · 计算机科学 2023-12-05 Tiancheng Qin , S. Rasoul Etesami

We propose a policy iteration method to solve an inverse problem for a mean-field game (MFG) model, specifically to reconstruct the obstacle function in the game from the partial observation data of value functions, which represent the…

最优化与控制 · 数学 2026-02-12 Kui Ren , Nathan Soedjak , Shanyin Tong

We consider infinite duration alternating move games. These games were previously studied by Roth, Balcan, Kalai and Mansour. They presented an FPTAS for computing an approximated equilibrium, and conjectured that there is a polynomial…

计算机科学与博弈论 · 计算机科学 2013-04-25 Yaron Velner

Stochastic games generalize Markov decision processes (MDPs) to a multiagent setting by allowing the state transitions to depend jointly on all player actions, and having rewards determined by multiplayer matrix games at each state. We…

计算机科学与博弈论 · 计算机科学 2013-01-18 Michael Kearns , Yishay Mansour , Satinder Singh

We study model-based and model-free policy optimization in a class of nonzero-sum stochastic dynamic games called linear quadratic (LQ) deep structured games. In such games, players interact with each other through a set of weighted…

计算机科学与博弈论 · 计算机科学 2020-12-15 Masoud Roudneshin , Jalal Arabneydi , Amir G. Aghdam

A zero-sum two person Perfect Information Stochastic game (PISG) under limiting average payoff has a value and both the maximiser and the minimiser have optimal pure stationary strategies. Firstly we form the matrix of undiscounted payoffs…

最优化与控制 · 数学 2023-02-15 K. G. Bakshi , S. Sinha

This paper introduces alignment games, a new class of zero-sum games modeling strategic interventions where effectiveness depends on alignment with an underlying hidden state. Motivated by operational problems in medical diagnostics,…

最优化与控制 · 数学 2025-09-08 Pedro Cesar Lopes Gerum , Thomas Lidbetter

We consider time-homogeneous uniformly nondegenerate stochastic differential games in domains and propose constructing $\varepsilon$-optimal strategies and policies by using adjoint Markov strategies and adjoint Markov policies which are…

最优化与控制 · 数学 2019-03-26 N. V. Krylov

Originating in evolutionary game theory, the class of "zero-determinant" strategies enables a player to unilaterally enforce linear payoff relationships in simple repeated games. An upshot of this kind of payoff constraint is that it can…

理论经济学 · 经济学 2025-11-26 Nikos Dimou , Alex McAvoy

This paper introduces an explicit algorithm for computing perfect public equilibrium (PPE) payoffs in repeated games with imperfect public monitoring, public randomization, and discounting. The method adapts the established framework by…

理论经济学 · 经济学 2024-11-05 Jasmina Karabegovic

Games, in their mathematical sense, are everywhere (game industries, economics, defense, education, chemistry, biology, ...).Search algorithms in games are artificial intelligence methods for playing such games. Unfortunately, there is no…

人工智能 · 计算机科学 2025-05-16 Quentin Cohen-Solal

We motivate and propose a new model for non-cooperative Markov game which considers the interactions of risk-aware players. This model characterizes the time-consistent dynamic "risk" from both stochastic state transitions (inherent to the…

计算机科学与博弈论 · 计算机科学 2019-11-22 Wenjie Huang , Pham Viet Hai , William B. Haskell

We suggest a new algorithm for two-person zero-sum undiscounted stochastic games focusing on stationary strategies. Given a positive real $\epsilon$, let us call a stochastic game $\epsilon$-ergodic, if its values from any two initial…

计算机科学与博弈论 · 计算机科学 2015-08-17 Endre Boros , Khaled Elbassioni , Vladimir Gurvich , Kazuhisa Makino

We consider zero sum stochastic games. For every discount factor $\lambda$, a time normalization allows to represent the game as being played on the interval [0, 1]. We introduce the trajectories of cumulated expected payoff and of…

最优化与控制 · 数学 2018-12-21 Sylvain Sorin , Guillaume Vigeral