Related papers: A Pseudo-Polynomial Algorithm for Mean Payoff Stoc…
In this paper, we consider a zero-sum undiscounted stochastic game which has finite state space and finitely many pure actions. Also, we assume the transition probability of the undiscounted stochastic game is controlled by one player and…
We study agents competing against each other in a repeated network zero-sum game while applying the multiplicative weights update (MWU) algorithm with fixed learning rates. In our implementation, agents select their strategies…
In this paper, we solve the constant-payoff conjecture formulated by Sorin, Venel and Vigeral (2010), for absorbing games with an arbitrary evaluation of the stage rewards. That is, the existence of a pair of asymptotically optimal…
Recently, Sidford, Wang, Wu and Ye (2018) developed an algorithm combining variance reduction techniques with value iteration to solve discounted Markov decision processes. This algorithm has a sublinear complexity when the discount factor…
We consider the problem of scheduling in multi-class, parallel-server queuing systems with uncertain rewards from job-server assignments. In this scenario, jobs incur holding costs while awaiting completion, and job-server assignments yield…
We study online reinforcement learning in average-reward stochastic games (SGs). An SG models a two-player zero-sum game in a Markov environment, where state transitions and one-step payoffs are determined simultaneously by a learner and an…
This paper considers a two-person zero-sum continuous-time Markov pure jump game in Borel state and action spaces over a fixed finite horizon. The main assumption on the model is the existence of a drift function, which bounds the reward…
A strategy profile in a multi-player game is a Nash equilibrium if no player can unilaterally deviate to achieve a strictly better payoff. A profile is an $\epsilon$-Nash equilibrium if no player can gain more than $\epsilon$ by…
We study optimal equilibria in multi-player games. An equilibrium is optimal for a player, if her payoff is maximal. A tempting approach to solving this problem is to seek optimal Nash equilibria, the standard form of equilibria where no…
We study a stochastic variant of the Team Orienteering Problem with lognormal travel times and an all-or-nothing reward policy, under which the reward of a route is lost if its travel time exceeds the available budget. We propose a…
We study two-player zero-sum stochastic games, and propose a form of independent learning dynamics called Doubly Smoothed Best-Response dynamics, which integrates a discrete and doubly smoothed variant of the best-response dynamics into…
Pursuit-evasion scenarios appear widely in robotics, security domains, and many other real-world situations. We focus on two-player pursuit-evasion games with concurrent moves, infinite horizon, and discounted rewards. We assume that the…
Classical objectives in two-player zero-sum games played on graphs often deal with limit behaviors of infinite plays: e.g., mean-payoff and total-payoff in the quantitative setting, or parity in the qualitative one (a canonical way to…
We consider turn-based stochastic two-player games with a combination of a parity condition that must hold surely, that is in all possible outcomes, and of a parity condition that must hold almost-surely, that is with probability 1. The…
An ever-important issue is protecting infrastructure and other valuable targets from a range of threats from vandalism to theft to piracy to terrorism. The "defender" can rarely afford the needed resources for a 100% protection. Thus, the…
This paper develops a unified framework for zero-sum games in which both the pure strategies and the payoff matrices contain complex-valued entries. By leveraging a linear isomorphism between complex and real vector spaces, we extend key…
We study multi-player games with perfect information and general payoff function, where the set of stages is the set of non-positive integers $\{\ldots,-2,-1,0\}$. We define two related equilibrium concepts: one considering only deviations…
We consider a zero-sum continuous time stopping game in which the pay-off is revealed in the maximum of the two stopping times instead of the minimum, which is the case in Dynkin games.
The paper investigates the long-time behavior of zero-sum linear-quadratic stochastic differential games, aiming to demonstrate that, under appropriate conditions, both the saddle strategy and the optimal state process exhibit the…
We consider infinite-state turn-based stochastic games of two players, Box and Diamond, who aim at maximizing and minimizing the expected total reward accumulated along a run, respectively. Since the total accumulated reward is unbounded,…