中文
相关论文

相关论文: Policy iteration algorithm for zero-sum multichain…

200 篇论文

We consider the discrete-time infinite-horizon optimal control problem formalized by Markov Decision Processes. We revisit the work of Bertsekas and Ioffe, that introduced $\lambda$ Policy Iteration, a family of algorithms parameterized by…

人工智能 · 计算机科学 2015-03-13 Bruno Scherrer

In this paper, we propose a new efficient algorithm to compute the value function for zero-sum stopping games featuring two players with opposing interests. This can be seen as a game version of the ''forward algorithm'' for (one-player)…

概率论 · 数学 2026-02-03 Nhat-Thang Le

It is well-known that for infinitely repeated games, there are computable strategies that have best responses, but no computable best responses. These results were originally proved for either specific games (e.g., Prisoner's dilemma), or…

计算机科学与博弈论 · 计算机科学 2020-06-11 Jakub Dargaj , Jakob Grue Simonsen

We study zero-sum games in the space of probability distributions over the Euclidean space $\mathbb{R}^d$ with entropy regularization, in the setting when the interaction function between the players is smooth and strongly convex-strongly…

计算机科学与博弈论 · 计算机科学 2025-07-01 Yang Cai , Siddharth Mitra , Xiuyuan Wang , Andre Wibisono

In this paper, we investigate a partially observable zero sum games where the state process is a discrete time Markov chain. We consider a general utility function in the optimization criterion. We show the existence of value for both…

最优化与控制 · 数学 2022-11-16 Arnab Bhabak , Subhamay saha

We study a two-player, zero-sum, stochastic game with incomplete information on one side in which the players are allowed to play more and more frequently. The informed player observes the realization of a Markov chain on which the payoffs…

最优化与控制 · 数学 2013-07-15 Pierre Cardaliaguet , Catherine Rainer , Dinah Rosenberg , Nicolas Vieille

We consider a zero-sum stochastic differential game over elementary mixed feed-back strategies. These are strategies based only on the knowledge of the past state, randomized continuously in time from a sampling distribution which is kept…

最优化与控制 · 数学 2014-04-16 Mihai Sîrbu

In this paper, we investigate a competitive market involving two agents who consider both their own wealth and the wealth gap with their opponent. Both agents can invest in a financial market consisting of a risk-free asset and a risky…

最优化与控制 · 数学 2025-02-10 Junyi Guo , Xia Han , Hao Wang , Kam Chuen Yuen

We consider both finite-state game graphs and recursive game graphs (or pushdown game graphs), that can model the control flow of sequential programs with recursion, with multi-dimensional mean-payoff objectives. In pushdown games two types…

计算机科学与博弈论 · 计算机科学 2013-08-09 Krishnendu Chatterjee , Yaron Velner

Two standard algorithms for approximately solving two-player zero-sum concurrent reachability games are value iteration and strategy iteration. We prove upper and lower bounds of 2^(m^(Theta(N))) on the worst case number of iterations…

计算机科学与博弈论 · 计算机科学 2012-03-02 Kristoffer Arnsfelt Hansen , Rasmus Ibsen-Jensen , Peter Bro Miltersen

Multi-dimensional mean-payoff and energy games provide the mathematical foundation for the quantitative study of reactive systems, and play a central role in the emerging quantitative theory of verification and synthesis. In this work, we…

计算机科学与博弈论 · 计算机科学 2014-11-04 Krishnendu Chatterjee , Mickael Randour , Jean-François Raskin

We propose a generic mechanism for incentivizing behavior in an arbitrary finite game using payments. Doing so is trivial if the mechanism is allowed to observe all actions taken in the game, as this allows it to simply punish those agents…

计算机科学与博弈论 · 计算机科学 2023-04-05 Nikolaj I. Schwartzbach

We address two-player general-sum stochastic Stackelberg games (SSGs), where the leader's policy is optimized considering the best-response follower whose policy is optimal for its reward under the leader. Existing policy gradient and value…

计算机科学与博弈论 · 计算机科学 2026-03-17 Mikoto Kudo , Youhei Akimoto

We consider multiplayer stochastic games in which the payoff of each player is a bounded and Borel-measurable function of the infinite play. By using a generalization of the technique of Martin (1998) and Maitra and Sudderth (1998), we show…

最优化与控制 · 数学 2022-08-26 János Flesch , Eilon Solan

Similar to the role of Markov decision processes in reinforcement learning, Stochastic Games (SGs) lay the foundation for the study of multi-agent reinforcement learning (MARL) and sequential agent interactions. In this paper, we derive…

计算机科学与博弈论 · 计算机科学 2023-01-12 Xiaotie Deng , Ningyuan Li , David Mguni , Jun Wang , Yaodong Yang

We consider stochastic control models with Borel spaces and universally measurable policies. For such models the standard policy iteration is known to have difficult measurability issues and cannot be carried out in general. We present a…

最优化与控制 · 数学 2016-02-26 Huizhen Yu , Dimitri P. Bertsekas

Zero-determinant strategies are a class of strategies in repeated games which unilaterally control payoffs. Zero-determinant strategies have attracted much attention in studies of social dilemma, particularly in the context of evolution of…

统计力学 · 物理学 2024-11-11 Masahiko Ueda

This paper proposes a new method for finding closed-loop saddle points in zero-sum linear-quadratic stochastic differential games by decoupling their inherent structure. Specifically, we develop a nested iterative scheme that constructs a…

最优化与控制 · 数学 2025-12-10 Yiyuan Wang

Computational advertising has been studied to design efficient marketing strategies that maximize the number of acquired customers. In an increased competitive market, however, a market leader (a leader) requires the acquisition of new…

计算机科学与博弈论 · 计算机科学 2019-06-18 Daisuke Hatano , Yuko Kuroki , Yasushi Kawase , Hanna Sumita , Naonori Kakimura , Ken-ichi Kawarabayashi

We study a class of stochastic target games where one player tries to find a strategy such that the state process almost-surely reaches a given target, no matter which action is chosen by the opponent. Our main result is a geometric dynamic…

概率论 · 数学 2015-02-03 Bruno Bouchard , Marcel Nutz