中文
相关论文

相关论文: On Bellman's Optimality Principle for zs-POSGs

200 篇论文

Stochastic games generalize Markov decision processes (MDPs) to a multiagent setting by allowing the state transitions to depend jointly on all player actions, and having rewards determined by multiplayer matrix games at each state. We…

计算机科学与博弈论 · 计算机科学 2013-01-18 Michael Kearns , Yishay Mansour , Satinder Singh

In this paper, we investigate a partially observable zero sum games where the state process is a discrete time Markov chain. We consider a general utility function in the optimization criterion. We show the existence of value for both…

最优化与控制 · 数学 2022-11-16 Arnab Bhabak , Subhamay saha

In this paper we study zero-sum two-player stochastic differential games with the help of theory of Backward Stochastic Differential Equations (BSDEs). At the one hand we generalize the results of the pioneer work of Fleming and Souganidis…

概率论 · 数学 2011-02-19 Rainer Buckdahn , Juan Li

It is well known that for any finite state Markov decision process (MDP) there is a memoryless deterministic policy that maximizes the expected reward. For partially observable Markov decision processes (POMDPs), optimal memoryless policies…

最优化与控制 · 数学 2016-02-16 Guido Montufar , Keyan Ghazi-Zahedi , Nihat Ay

In the present paper, we study a two-player zero-sum deterministic differential game with both players adopting impulse controls, in infinite time horizon, under rather weak assumptions on the cost functions. We prove by means of the…

最优化与控制 · 数学 2021-01-29 Brahim El Asri , Hafid Lalioui , Sehail Mazid

We extend the construction of equilibria for linear-quadratic and mean-variance portfolio problems available in the literature to a large class of mean-field time-inconsistent stochastic control problems in continuous time. Our approach…

最优化与控制 · 数学 2021-10-01 Jiang Yu Nguwi , Nicolas Privault

We study two-player concurrent stochastic games on finite graphs, with B\"uchi and co-B\"uchi objectives. The goal of the first player is to maximize the probability of satisfying the given objective. Following Martin's determinacy theorem…

计算机科学与博弈论 · 计算机科学 2022-11-28 Benjamin Bordais , Patricia Bouyer , Stéphane Le Roux

This paper proposes a new mathematical paradigm to analyze discrete-time mean-field games. It is shown that finding Nash equilibrium solutions for a general class of discrete-time mean-field games is equivalent to solving an optimization…

最优化与控制 · 数学 2023-08-29 Xin Guo , Anran Hu , Junzi Zhang

This paper is concerned with the axiomatic foundation and explicit construction of a general class of optimality criteria that can be used for investment problems with multiple time horizons, or when the time horizon is not known in…

投资组合管理 · 定量金融 2014-02-03 Sergey Nadtochiy , Michael Tehranchi

Traditional mean-field game (MFG) solvers operate on an instance-by-instance basis, which becomes infeasible when many related problems must be solved (e.g., for seeking a robust description of the solution under perturbations of the…

最优化与控制 · 数学 2025-10-24 Dena Firoozi , Anastasis Kratsios , Xuwei Yang

This paper analyzes a class of infinite-time-horizon stochastic games with singular controls motivated from the partially reversible problem. It provides an explicit solution for the mean-field game (MFG) and presents sensitivity analysis…

最优化与控制 · 数学 2020-08-12 Haoyang Cao , Xin Guo

We develop provably efficient reinforcement learning algorithms for two-player zero-sum finite-horizon Markov games with simultaneous moves. To incorporate function approximation, we consider a family of Markov games where the reward…

机器学习 · 计算机科学 2020-06-25 Qiaomin Xie , Yudong Chen , Zhaoran Wang , Zhuoran Yang

The objective of this work is to study continuous-time Markov decision processes on a general Borel state space with both impulsive and continuous controls for the infinite-time horizon discounted cost. The continuous-time controlled…

最优化与控制 · 数学 2019-08-17 François Dufour , Alexei Piunovskiy

We provide a general approach to reformulating any continuous-time stochastic Stackelberg differential game under closed-loop strategies as a single-level optimisation problem with target constraints. More precisely, we consider a…

最优化与控制 · 数学 2026-05-14 Camilo Hernández , Nicolás Hernández Santibáñez , Emma Hubert , Dylan Possamaï

Motivated by empirical evidence that individuals within group decision making simultaneously aspire to maximize utility and avoid inequality we propose a criterion based on the entropy-norm pair for geometric selection of strict Nash…

物理与社会 · 物理学 2020-03-23 A. B. Leoneti , G. A. Prataviera

We consider a variant of the hide-and-seek game in which a seeker inspects multiple hiding locations to find multiple items hidden by a hider. Each hiding location has a maximum hiding capacity and a probability of detecting its hidden…

计算机科学与博弈论 · 计算机科学 2024-06-26 Bastián Bahamondes , Mathieu Dahan

Network congestion games are a convenient model for reasoning about routing problems in a network: agents have to move from a source to a target vertex while avoiding congestion, measured as a cost depending on the number of players using…

计算机科学与博弈论 · 计算机科学 2022-07-05 Aline Goeminne , Nicolas Markey , Ocan Sankur

We consider a two-player zero-sum deterministic differential game where each player uses both continuous and impulse controls in infinite-time horizon. We assume that the impulses supposed to be of general term and the costs depend on the…

最优化与控制 · 数学 2022-09-26 Brahim El Asri , Hafid Lalioui

The use of pessimism, when reasoning about datasets lacking exhaustive exploration has recently gained prominence in offline reinforcement learning. Despite the robustness it adds to the algorithm, overly pessimistic reasoning can be…

机器学习 · 计算机科学 2023-10-25 Tengyang Xie , Ching-An Cheng , Nan Jiang , Paul Mineiro , Alekh Agarwal

This paper deals with the unconstrained and constrained cases for continuous-time Markov decision processes under the finite-horizon expected total cost criterion. The state space is denumerable and the transition and cost rates are allowed…

最优化与控制 · 数学 2014-08-26 Qingda Wei , Xian Chen