中文
相关论文

相关论文: Expected Window Mean-Payoff

200 篇论文

The paper is concerned with the dependence of the solution of the deterministic mean field game on the initial distribution of players. The main object of study is the mapping which assigns to the initial time and the initial distribution…

最优化与控制 · 数学 2019-02-27 Yurii Averboukh

In a two-player zero-sum graph game the players move a token throughout a graph to produce an infinite path, which determines the winner or payoff of the game. Traditionally, the players alternate turns in moving the token. In {\em bidding…

理论经济学 · 经济学 2020-12-22 Guy Avni , Ismaël Jecker , Đorđe Žikelić

Graph games provide the foundation for modeling and synthesizing reactive processes. In the synthesis of stochastic reactive processes, the traditional model is perfect-information stochastic games, where some transitions of the game graph…

计算机科学中的逻辑 · 计算机科学 2016-04-22 Krishnendu Chatterjee , Laurent Doyen

In this paper, we study the notion of adversarial Stackelberg value for two-player non-zero sum games played on bi-weighted graphs with the mean-payoff and the discounted sum functions. The adversarial Stackelberg value of Player 0 is the…

计算机科学与博弈论 · 计算机科学 2020-05-05 Emmanuel Filiot , Raffaella Gentilini , Jean-François Raskin

We consider both finite-state game graphs and recursive game graphs (or pushdown game graphs), that can model the control flow of sequential programs with recursion, with multi-dimensional mean-payoff objectives. In pushdown games two types…

计算机科学与博弈论 · 计算机科学 2013-08-09 Krishnendu Chatterjee , Yaron Velner

Mean-payoff games play a central role in quantitative synthesis and verification. In a single-dimensional game a weight is assigned to every transition and the objective of the protagonist is to assure a non-negative limit-average weight.…

计算机科学中的逻辑 · 计算机科学 2014-10-22 Yaron Velner

We consider the batch (off-line) policy learning problem in the infinite horizon Markov Decision Process. Motivated by mobile health applications, we focus on learning a policy that maximizes the long-term average reward. We propose a…

统计理论 · 数学 2022-09-20 Peng Liao , Zhengling Qi , Runzhe Wan , Predrag Klasnja , Susan Murphy

We consider partially observable Markov decision processes (POMDPs) with limit-average payoff, where a reward value in the interval [0,1] is associated to every transition, and the payoff of an infinite path is the long-run average of the…

人工智能 · 计算机科学 2014-08-12 Krishnendu Chatterjee , Martin Chmelik

We consider mean-field control problems in discrete time with discounted reward, infinite time horizon and compact state and action space. The existence of optimal policies is shown and the limiting mean-field problem is derived when the…

最优化与控制 · 数学 2025-10-16 Nicole Bäuerle

While discounted payoff games and classic games that reduce to them, like parity and mean-payoff games, are symmetric, their solutions are not. We have taken a fresh view on the properties that optimal solutions need to have, and devised a…

数据结构与算法 · 计算机科学 2026-03-11 Daniele Dell'Erba , Arthur Dumas , Sven Schewe

In this paper, we provide an effective characterization of all the subgame-perfect equilibria in infinite duration games played on finite graphs with mean-payoff objectives. To this end, we introduce the notion of requirement, and the…

计算机科学与博弈论 · 计算机科学 2024-02-14 Léonard Brice , Marie van den Bogaard , Jean-François Raskin

An average-time game is played on the infinite graph of configurations of a finite timed automaton. The two players, Min and Max, construct an infinite run of the automaton by taking turns to perform a timed transition. Player Min wants to…

计算机科学与博弈论 · 计算机科学 2020-01-16 Marcin Jurdzinski , Ashutosh Trivedi

While discounted payoff games and classic games that reduce to them, like parity and mean-payoff games, are symmetric, their solutions are not. We have taken a fresh view on the constraints that optimal solutions need to satisfy, and…

数据结构与算法 · 计算机科学 2023-10-03 Daniele Dell'Erba , Arthur Dumas , Sven Schewe

Markov decision processes (MDPs) are standard models for probabilistic systems with non-deterministic behaviours. Mean payoff (or long-run average reward) provides a mathematically elegant formalism to express performance related…

性能 · 计算机科学 2017-09-08 Jan Křetínský , Tobias Meggendorfer

We consider concurrent games played on graphs. At every round of the game, each player simultaneously and independently selects a move; the moves jointly determine the transition to a successor state. Two basic objectives are the safety…

计算机科学与博弈论 · 计算机科学 2008-12-18 Krishnendu Chatterjee , Luca de Alfaro , Thomas A. Henzinger

Consider a very simple class of (finite) games: after an initial move by nature, each player makes one move. Moreover, the players have common interests: at each node, all the players get the same payoff. We show that the problem of…

计算机科学与博弈论 · 计算机科学 2007-05-23 Francis Chu , Joseph Y. Halpern

A valuation for a player in a game in extensive form is an assignment of numeric values to the players moves. The valuation reflects the desirability moves. We assume a myopic player, who chooses a move with the highest valuation.…

机器学习 · 计算机科学 2007-05-23 Philippe Jehiel , Dov Samet

Evolutionary game theory is a powerful mathematical framework to study how intelligent individuals adjust their strategies in collective interactions. It has been widely believed that it is impossible to unilaterally control players'…

最优化与控制 · 数学 2021-08-31 Renfei Tan , Qi Su , Bin Wu , Long Wang

We study the complexity of central controller synthesis problems for finite-state Markov decision processes, where the objective is to optimize both the expected mean-payoff performance of the system and its stability. We argue that the…

系统与控制 · 计算机科学 2013-05-20 Tomáš Brázdil , Krishnendu Chatterjee , Vojtěch Forejt , Antonín Kučera

We consider finite horizon Markov decision processes under performance measures that involve both the mean and the variance of the cumulative reward. We show that either randomized or history-based policies can improve performance. We prove…

机器学习 · 计算机科学 2011-05-02 Shie Mannor , John Tsitsiklis