中文
相关论文

相关论文: Expected Window Mean-Payoff

200 篇论文

The Shapley value equals a player's contribution to the potential of a game. The potential is a most natural one-number summary of a game, which can be computed as the expected accumulated worth of a random partition of the players. This…

理论经济学 · 经济学 2024-07-23 André Casajus , Yukihiko Funaki , Frank Huettner

We consider Markov decision processes where the state of the chain is only given at chosen observation times and of a cost. Optimal strategies involve the optimisation of observation times as well as the subsequent action values. We…

最优化与控制 · 数学 2025-03-27 Christoph Reisinger , Jonathan Tam

Semidefinite programming can be considered over any real closed field, including fields of Puiseux series equipped with their nonarchimedean valuation. Nonarchimedean semidefinite programs encode parametric families of classical…

最优化与控制 · 数学 2018-02-22 Xavier Allamigeon , Stéphane Gaubert , Ricardo D. Katz , Mateusz Skomra

A classic solution technique for Markov decision processes (MDP) and stochastic games (SG) is value iteration (VI). Due to its good practical performance, this approximative approach is typically preferred over exact techniques, even though…

人工智能 · 计算机科学 2023-04-21 Jan Křetínský , Tobias Meggendorfer , Maximilian Weininger

In this paper, we consider the problem of optimization and learning for constrained and multi-objective Markov decision processes, for both discounted rewards and expected average rewards. We formulate the problems as zero-sum games where…

最优化与控制 · 数学 2021-03-05 Ather Gattami , Qinbo Bai , Vaneet Agarwal

We design and analyze minimax-optimal algorithms for online linear optimization games where the player's choice is unconstrained. The player strives to minimize regret, the difference between his loss and the loss of a post-hoc benchmark…

机器学习 · 计算机科学 2013-02-12 H. Brendan McMahan

We consider a nonzero-sum N-player Markov game on an abstract measurable state space with compact metric action spaces. The payoff functions are bounded Carath\'eodory functions and the transitions of the system are assumed to have a…

最优化与控制 · 数学 2023-05-09 François Dufour , Tomás Prieto-Rumeau

In game theory and multi-agent reinforcement learning (MARL), each agent selects a strategy, interacts with the environment and other agents, and subsequently updates its strategy based on the received payoff. This process generates a…

计算机科学与博弈论 · 计算机科学 2025-09-30 Yanqing Fu , Chao Huang , Chenrun Wang , Zhuping Wang

The interval scheduling problem is one variant of the scheduling problem. In this paper, we propose a novel variant of the interval scheduling problem, whose definition is as follows: given jobs are specified by their {\em release times},…

数据结构与算法 · 计算机科学 2018-05-16 Koji M. Kobayashi

We consider multiple-environment Markov decision processes (MEMDP), which consist of a finite set of MDPs over the same state space, representing different scenarios of transition structure and probability. The value of a strategy is the…

计算机科学中的逻辑 · 计算机科学 2025-04-23 Krishnendu Chatterjee , Laurent Doyen , Jean-François Raskin , Ocan Sankur

In a probabilistic mean field game driven by a L\'evy process an individual player aims to minimize a long run discounted/ergodic cost by controlling the process through a pair of increasing and decreasing c\`adl\`ag processes, while he is…

最优化与控制 · 数学 2025-05-30 Facundo Oliú

Energy Markov Decision Processes (EMDPs) are finite-state Markov decision processes where each transition is assigned an integer counter update and a rational payoff. An EMDP configuration is a pair s(n), where s is a control state and n is…

计算机科学中的逻辑 · 计算机科学 2016-07-05 Tomáš Brázdil , Antonín Kučera , Petr Novotný

This paper analyzes a simple game with $n$ players. We fix a mean, $\mu$, in the interval $[0, 1]$ and let each player choose any random variable distributed on that interval with the given mean. The winner of the zero-sum game is the…

概率论 · 数学 2018-04-24 Artem Hulko , Mark Whitmeyer

Graph games are fundamental in strategic reasoning of multi-agent systems and their environments. We study a new family of graph games which combine stochastic environmental uncertainties and auction-based interactions among the agents,…

计算机科学与博弈论 · 计算机科学 2024-12-30 Guy Avni , Martin Kurečka , Kaushik Mallik , Petr Novotný , Suman Sadhukhan

The paper is concerned with a zero-sum continuous-time stochastic differential game with a dynamics controlled by a Markov process and a terminal payoff. The value function of the original game is estimated using the value function of a…

最优化与控制 · 数学 2016-02-16 Yurii Averboukh

We consider infinite duration alternating move games. These games were previously studied by Roth, Balcan, Kalai and Mansour. They presented an FPTAS for computing an approximated equilibrium, and conjectured that there is a polynomial…

计算机科学与博弈论 · 计算机科学 2013-04-25 Yaron Velner

Mechanism design is a well-established game-theoretic paradigm for designing games to achieve desired outcomes. This paper addresses a closely related but distinct concept, equilibrium design. Unlike mechanism design, the designer's…

计算机科学与博弈论 · 计算机科学 2024-08-20 Muhammad Najib , Giuseppe Perelli

We study the price of anarchy in a class of graph coloring games (a subclass of polymatrix common-payoff games). In those games, players are vertices of an undirected, simple graph, and the strategy space of each player is the set of colors…

计算机科学与博弈论 · 计算机科学 2015-10-05 Lasse Kliemann , Elmira Shirazi Sheykhdarabadi , Anand Srivastav

Value iteration is a fundamental algorithm for solving Markov Decision Processes (MDPs). It computes the maximal $n$-step payoff by iterating $n$ times a recurrence equation which is naturally associated to the MDP. At the same time, value…

形式语言与自动机理论 · 计算机科学 2019-04-30 Nikhil Balaji , Stefan Kiefer , Petr Novotný , Guillermo A. Pérez , Mahsa Shirmohammadi

We propose a new complexity measure for Markov decision processes (MDPs), the maximum expected hitting cost (MEHC). This measure tightens the closely related notion of diameter [JOA10] by accounting for the reward structure. We show that…

机器学习 · 计算机科学 2019-11-06 Falcon Z. Dai , Matthew R. Walter
‹ 上一页 1 8 9 10 下一页 ›