中文
相关论文

相关论文: Policy iteration for perfect information stochasti…

200 篇论文

We devise a policy-iteration algorithm for deterministic two-player discounted and mean-payoff games, that runs in polynomial time with high probability, on any input where each payoff is chosen independently from a sufficiently random…

计算机科学与博弈论 · 计算机科学 2024-02-07 Bruno Loff , Mateusz Skomra

We consider zero-sum stochastic games with finite state and action spaces, perfect information, mean payoff criteria, without any irreducibility assumption on the Markov chains associated to strategies (multichain games). The value of such…

最优化与控制 · 数学 2012-08-03 Marianne Akian , Jean Cochet-Terrasson , Sylvie Detournay , Stéphane Gaubert

Ye showed recently that the simplex method with Dantzig pivoting rule, as well as Howard's policy iteration algorithm, solve discounted Markov decision processes (MDPs), with a constant discount factor, in strongly polynomial time. More…

计算机科学与博弈论 · 计算机科学 2010-08-04 Thomas Dueholm Hansen , Peter Bro Miltersen , Uri Zwick

We examine perfect information stochastic mean-payoff games - a class of games containing as special sub-classes the usual mean-payoff games and parity games. We show that deterministic memoryless strategies that are optimal for discounted…

计算机科学与博弈论 · 计算机科学 2010-06-09 Hugo Gimbert , Wiesław Zielonka

This note provides upper bounds on the number of operations required to compute by value iterations a nearly optimal policy for an infinite-horizon discounted Markov decision process with a finite number of states and actions. For a given…

最优化与控制 · 数学 2020-01-29 Eugene A. Feinberg , Gaojin He

Two-player zero-sum repeated games are well understood. Computing the value of such a game is straightforward. Additionally, if the payoffs are dependent on a random state of the game known to one, both, or neither of the players, the…

信息论 · 计算机科学 2009-11-05 Paul Cuff

Stochastic games are an important class of problems that generalize Markov decision processes to game theoretic scenarios. We consider finite state two-player zero-sum stochastic games over an infinite time horizon with discounted rewards.…

最优化与控制 · 数学 2008-06-17 Parikshit Shah , Pablo A. Parrilo

We examine the problem of the existence of optimal deterministic stationary strategiesintwo-players antagonistic (zero-sum) perfect information stochastic games with finitely many states and actions.We show that the existenceof such…

计算机科学与博弈论 · 计算机科学 2016-11-28 Hugo Gimbert , Wieslaw Zielonka

Policy-based methods with function approximation are widely used for solving two-player zero-sum games with large state and/or action spaces. However, it remains elusive how to obtain optimization and statistical guarantees for such…

机器学习 · 计算机科学 2022-03-01 Yulai Zhao , Yuandong Tian , Jason D. Lee , Simon S. Du

This work considers two-player zero-sum semi-Markov games with incomplete information on one side and perfect observation. At the beginning, the system selects a game type according to a given probability distribution and informs to Player…

最优化与控制 · 数学 2021-07-16 Fang Chen , Xianping Guo , Zhong-Wei Liao

We consider some well-known families of two-player, zero-sum, perfect information games that can be viewed as special cases of Shapley's stochastic games. We show that the following tasks are polynomial time equivalent: - Solving simple…

计算机科学与博弈论 · 计算机科学 2008-12-03 Vladimir Gurvich , Peter Bro Miltersen

We study zero-sum differential games with state constraints and one-sided information, where the informed player (Player 1) has a categorical payoff type unknown to the uninformed player (Player 2). The goal of Player 1 is to minimize his…

计算机科学与博弈论 · 计算机科学 2024-06-05 Mukesh Ghimire , Lei Zhang , Zhe Xu , Yi Ren

Graph games provide the foundation for modeling and synthesizing reactive processes. In the synthesis of stochastic reactive processes, the traditional model is perfect-information stochastic games, where some transitions of the game graph…

计算机科学中的逻辑 · 计算机科学 2016-04-22 Krishnendu Chatterjee , Laurent Doyen

The value of a finite-state two-player zero-sum stochastic game with limit-average payoff can be approximated to within $\epsilon$ in time exponential in a polynomial in the size of the game times polynomial in logarithmic in…

计算机科学与博弈论 · 计算机科学 2008-12-18 Krishnendu Chatterjee , Rupak Majumdar , Thomas A. Henzinger

We consider concurrent mean-payoff games, a very well-studied class of two-player (player 1 vs player 2) zero-sum games on finite-state graphs where every transition is assigned a reward between 0 and 1, and the payoff function is the…

计算机科学与博弈论 · 计算机科学 2014-10-02 Krishnendu Chatterjee , Rasmus Ibsen-Jensen

We present a polynomial-time reduction from max-plus-average constraints to the feasibility problem for semidefinite programs. This shows that Condon's simple stochastic games, stochastic mean payoff games, and in particular mean payoff…

最优化与控制 · 数学 2025-12-03 Manuel Bodirsky , Georg Loho , Mateusz Skomra

We introduce two-level discounted games played by two players on a perfect-information stochastic game graph. The upper level game is a discounted game and the lower level game is an undiscounted reachability game. Two-level games model…

计算机科学中的逻辑 · 计算机科学 2010-06-09 Krishnendu Chatterjee , Rupak Majumdar

We study policy iteration for infinite-horizon Markov decision processes. It has recently been shown policy iteration style algorithms have exponential lower bounds in a two player game setting. We extend these lower bounds to Markov…

数据结构与算法 · 计算机科学 2010-03-18 John Fearnley

We consider zero-sum stochastic games with perfect information and finitely many states and actions. The payoff is computed by a function which associates to each infinite sequence of states and actions a real number. We prove that if the…

计算机科学与博弈论 · 计算机科学 2022-03-29 Hugo Gimbert , Edon Kelmendi

A basic question for zero-sum repeated games consists in determining whether the mean payoff per time unit is independent of the initial state. In the special case of "zero-player" games, i.e., of Markov chains equipped with additive…

最优化与控制 · 数学 2015-10-20 Marianne Akian , Stéphane Gaubert , Antoine Hochart
‹ 上一页 1 2 3 10 下一页 ›