中文
相关论文

相关论文: On Bellman's Optimality Principle for zs-POSGs

200 篇论文

We first study an optimal stopping problem in which a player (an agent) uses a discrete stopping time in order to stop optimally a payoff process whose risk is evaluated by a (non-linear) $g$-expectation. We then consider a non-zero-sum…

概率论 · 数学 2017-05-11 Miryana Grigorova , Marie-Claire Quenez

We study optimality for the safety-constrained Markov decision process which is the underlying framework for safe reinforcement learning. Specifically, we consider a constrained Markov decision process (with finite states and finite…

系统与控制 · 电气工程与系统科学 2023-07-13 Rahul Misra , Rafał Wisniewski , Carsten Skovmose Kallesøe

This paper addresses the problem of learning a Nash equilibrium in $\gamma$-discounted multiplayer general-sum Markov Games (MG). A key component of this model is the possibility for the players to either collaborate or team apart to…

计算机科学与博弈论 · 计算机科学 2017-03-07 Julien Pérolat , Florian Strub , Bilal Piot , Olivier Pietquin

In this paper, I introduce a novel benchmark in games, super-Nash performance, and a solution concept, optimin, whereby players maximize their minimal payoff under unilateral profitable deviations by other players. Optimin achieves…

理论经济学 · 经济学 2025-10-23 Mehmet S. Ismail

The paper is concerned with a variant of the continuous-time finite state Markov game of control and stopping where both players can affect transition rates, while only one player can choose a stopping time. We use the dynamic programming…

最优化与控制 · 数学 2022-08-09 Yurii Averboukh

Zero-sum stochastic games provide a rich model for competitive decision making. However, under general forms of state uncertainty as considered in the Partially Observable Stochastic Game (POSG), such decision making problems are still not…

人工智能 · 计算机科学 2016-06-23 Auke J. Wiggers , Frans A. Oliehoek , Diederik M. Roijers

We consider the problem of optimally utilizing $N$ resources, each in an unknown binary state. The state of each resource can be inferred from state-dependent noisy measurements. Depending on its state, utilizing a resource results in…

系统与控制 · 计算机科学 2017-05-18 Lorenzo Ferrari , Qing Zhao , Anna Scaglione

This paper considers the problem of designing optimal algorithms for reinforcement learning in two-player zero-sum games. We focus on self-play algorithms which learn the optimal policy by playing against itself without any direct…

机器学习 · 计算机科学 2020-07-15 Yu Bai , Chi Jin , Tiancheng Yu

We consider the problem of finding the best memoryless stochastic policy for an infinite-horizon partially observable Markov decision process (POMDP) with finite state and action spaces with respect to either the discounted or mean reward…

最优化与控制 · 数学 2022-05-02 Johannes Müller , Guido Montúfar

We propose a new objective function for finite-horizon episodic Markov decision processes that better captures Bellman's principle of optimality, and provide an expression for the gradient of the objective.

机器学习 · 计算机科学 2019-06-10 Philip S. Thomas , Scott M. Jordan , Yash Chandak , Chris Nota , James Kostas

We contribute the first provable guarantees of global convergence to Nash equilibria (NE) in two-player zero-sum convex Markov games (cMGs) by using independent policy gradient methods. Convex Markov games, recently defined by Gemp et al.…

计算机科学与博弈论 · 计算机科学 2025-06-23 Fivos Kalogiannis , Emmanouil-Vasileios Vlatakis-Gkaragkounis , Ian Gemp , Georgios Piliouras

Computing approximate Nash equilibria in multi-player general-sum Markov games is a computationally intractable task. However, multi-player Markov games with certain cooperative or competitive structures might circumvent this…

计算机科学与博弈论 · 计算机科学 2023-08-17 Zailin Ma , Jiansheng Yang , Zhihua Zhang

The overall aim of our research is to develop techniques to reason about the equilibrium properties of multi-agent systems. We model multi-agent systems as concurrent games, in which each player is a process that is assumed to act…

计算机科学中的逻辑 · 计算机科学 2020-08-14 Julian Gutierrez , Aniello Murano , Giuseppe Perelli , Sasha Rubin , Thomas Steeples , Michael Wooldridge

This paper characterizes differentiable and subgame Markov perfect equilibria in a continuous time intertemporal decision problem with non-constant discounting. Capturing the idea of non commitment by letting the commitment period being…

最优化与控制 · 数学 2008-08-29 Ivar Ekeland , Ali Lazrak

We consider the problem of computing a mixed-strategy generalized Nash equilibrium (MS-GNE) for a class of games where each agent has both continuous and integer decision variables. Specifically, we propose a novel Bregman…

最优化与控制 · 数学 2022-06-14 Wicak Ananduta , Sergio Grammatico

Stochastic games are an important class of problems that generalize Markov decision processes to game theoretic scenarios. We consider finite state two-player zero-sum stochastic games over an infinite time horizon with discounted rewards.…

最优化与控制 · 数学 2008-06-17 Parikshit Shah , Pablo A. Parrilo

This paper studies partially observable two-person zero-sum semi-Markov games under a probability criterion, in which the system state may not be completely observed. It focuses on the probability that the accumulated rewards of player 1…

最优化与控制 · 数学 2025-08-26 Xin Wen , Li Xia , Zhihui Yu

This paper aims to solve the optimal strategy against a well-known adaptive algorithm, the Hedge algorithm, in a finitely repeated $2\times 2$ zero-sum game. In the literature, related theoretical results are very rare. To this end, we make…

最优化与控制 · 数学 2023-12-18 Xinxiang Guo , Yifen Mu

Many security and other real-world situations are dynamic in nature and can be modelled as strictly competitive (or zero-sum) dynamic games. In these domains, agents perform actions to affect the environment and receive observations --…

计算机科学与博弈论 · 计算机科学 2020-10-23 Karel Horák , Branislav Bošanský , Vojtěch Kovařík , Christopher Kiekintveld

In this paper, we consider discrete-time partially observed mean-field games with the risk-sensitive optimality criterion. We introduce risk-sensitivity behaviour for each agent via an exponential utility function. In the game model, each…

系统与控制 · 电气工程与系统科学 2022-11-11 Naci Saldi , Tamer Basar , Maxim Raginsky