中文
相关论文

相关论文: Solving Two-State Markov Games with Incomplete Inf…

200 篇论文

The famous theorem of R.Aumann and M.Maschler states that the sequence of values of an N-stage zero-sum game G_N with incomplete information on one side converges as N tends to infinity, and the error term is bounded by a constant divided…

计算机科学与博弈论 · 计算机科学 2013-12-30 Fedor Sandomirskiy

We propose a new framework of Markov $\alpha$-potential games to study Markov games. We show that any Markov game with finite-state and finite-action is a Markov $\alpha$-potential game, and establish the existence of an associated…

计算机科学与博弈论 · 计算机科学 2025-04-02 Xin Guo , Xinyu Li , Chinmay Maheshwari , Shankar Sastry , Manxi Wu

Strategy iteration is a technique frequently used for two-player games in order to determine the winner or compute payoffs, but to the best of our knowledge no general framework for strategy iteration has been considered. Inspired by…

计算机科学中的逻辑 · 计算机科学 2022-12-14 Paolo Baldan , Richard Eggert , Barbara König , Tommaso Padoan

We provide an algorithm to find the value and an optimal strategy of the solitaire variant of the Ten Thousand dice game in the framework of Markov Control Processes. Once an optimal critical threshold is found, the set of non-stopping…

最优化与控制 · 数学 2014-05-30 Fabián Crocce , Ernesto Mordecki

Reactive synthesis is a class of methods to construct a provably-correct control system, referred to as a robot, with respect to a temporal logic specification in the presence of a dynamic and uncontrollable environment. This is achieved by…

形式语言与自动机理论 · 计算机科学 2020-04-24 Abhishek N. Kulkarni , Jie Fu

For a two-player imperfect-information extensive-form game (IIEFG) with $K$ time steps and a player action space of size $U$, the game tree complexity is $U^{2K}$, causing existing IIEFG solvers to struggle with large or infinite $(U,K)$,…

计算机科学与博弈论 · 计算机科学 2026-03-03 Mukesh Ghimire , Lei Zhang , Zhe Xu , Yi Ren

We develop a probabilistic approach to continuous-time finite state mean field games. Based on an alternative description of continuous-time Markov chain by means of semimartingale and the weak formulation of stochastic optimal control, our…

概率论 · 数学 2018-08-24 Rene Carmona , Peiqi Wang

It is well-known that for infinitely repeated games, there are computable strategies that have best responses, but no computable best responses. These results were originally proved for either specific games (e.g., Prisoner's dilemma), or…

计算机科学与博弈论 · 计算机科学 2020-06-11 Jakub Dargaj , Jakob Grue Simonsen

Two-player games on graphs provide the theoretical frame- work for many important problems such as reactive synthesis. While the traditional study of two-player zero-sum games has been extended to multi-player games with several notions of…

计算机科学与博弈论 · 计算机科学 2013-11-14 Krishnendu Chatterjee , Laurent Doyen , Emmanuel Filiot , Jean-François Raskin

Policy gradient methods have become a staple of any single-agent reinforcement learning toolbox, due to their combination of desirable properties: iterate convergence, efficient use of stochastic trajectory feedback, and theoretically-sound…

计算机科学与博弈论 · 计算机科学 2025-07-10 Mingyang Liu , Gabriele Farina , Asuman Ozdaglar

This paper studies the optimization of strategies in the context of possibly randomized two players zero-sum games with incomplete information. We compare 5 algorithms for tuning the parameters of strategies over a benchmark of 12 games. A…

计算机科学与博弈论 · 计算机科学 2018-07-06 Marie-Liesse Cauwet , Olivier Teytaud

A general model for zero-sum stochastic games with asymmetric information is considered. In this model, each player's information at each time can be divided into a common information part and a private information part. Under certain…

系统与控制 · 电气工程与系统科学 2019-12-25 Dhruva Kartik , Ashutosh Nayyar

In a reachability-time game, players Min and Max choose moves so that the time to reach a final state in a timed automaton is minimised or maximised, respectively. Asarin and Maler showed decidability of reachability-time games on strongly…

计算复杂性 · 计算机科学 2020-01-16 Marcin Jurdziński , Ashutosh Trivedi

Secure equilibrium is a refinement of Nash equilibrium, which provides some security to the players against deviations when a player changes his strategy to another best response strategy. The concept of secure equilibrium is specifically…

计算机科学与博弈论 · 计算机科学 2014-05-08 Julie De Pril , János Flesch , Jeroen Kuipers , Gijs Schoenmakers , Koos Vrieze

We study a multi-agent reinforcement learning dynamics, and analyze its asymptotic behavior in infinite-horizon discounted Markov potential games. We focus on the independent and decentralized setting, where players do not know the game…

机器学习 · 计算机科学 2025-04-02 Chinmay Maheshwari , Manxi Wu , Druv Pai , Shankar Sastry

This report investigates the optimal design of event-triggered estimation for first-order linear stochastic systems. The problem is posed as a two-player team problem with a partially nested information pattern. The two players are given by…

最优化与控制 · 数学 2012-03-23 Adam Molin , Sandra Hirche

We consider a two-player game in which the first player (the Guesser) tries to guess, edge-by-edge, the path that second player (the Chooser) takes through a directed graph. At each step, the Guesser makes a wager as to the correctness of…

概率论 · 数学 2009-07-14 Marcus Pendergrass

Deep Reinforcement Learning combined with Fictitious Play shows impressive results on many benchmark games, most of which are, however, single-stage. In contrast, real-world decision making problems may consist of multiple stages, where the…

机器学习 · 计算机科学 2023-03-08 Wei Xi , Yongxin Zhang , Changnan Xiao , Xuefeng Huang , Shihong Deng , Haowei Liang , Jie Chen , Peng Sun

We study best-response type learning dynamics for zero-sum polymatrix games under two information settings. The two settings are distinguished by the type of information that each player has about the game and their opponents' strategy. The…

最优化与控制 · 数学 2025-08-13 Fathima Zarin Faizal , Asuman Ozdaglar , Martin J. Wainwright

This paper investigates the two-person zero-sum stochastic games for piece-wise deterministic Markov decision processes with risk-sensitive finite-horizon cost criterion on a general state space. Here, the transition and cost/reward rates…

最优化与控制 · 数学 2024-05-15 Subrata Golui