中文
相关论文

相关论文: Promises Made, Promises Kept: Safe Pareto Improvem…

200 篇论文

Graph games provide the foundation for modeling and synthesizing reactive processes. In the synthesis of stochastic reactive processes, the traditional model is perfect-information stochastic games, where some transitions of the game graph…

计算机科学中的逻辑 · 计算机科学 2016-04-22 Krishnendu Chatterjee , Laurent Doyen

One practical requirement in solving dynamic games is to ensure that the players play well from any decision point onward. To satisfy this requirement, existing efforts focus on equilibrium refinement, but the scalability and applicability…

多智能体系统 · 计算机科学 2021-08-24 Weizhe Chen , Zihan Zhou , Yi Wu , Fei Fang

Training models through self-play alone (without any human data) has been a longstanding goal in AI, but its effectiveness for training large language models remains unclear, particularly in code generation where rewards based on unit tests…

Game theoretic approaches have gained traction as robust methodologies for designing distributed local algorithms that induce a desired overall system configuration in multi-agent settings. However, much of the emphasis in these approaches…

系统与控制 · 电气工程与系统科学 2021-11-03 Rohit Konda , Rahul Chandan , David Grimsman , Jason R. Marden

We identify a subtle security issue that impacts mechanism design in scenarios in which agents can absolutely commit to strategies. Absolute commitments allow the strategy of an agent to depend on the commitments made by the other agents.…

计算机科学与博弈论 · 计算机科学 2024-01-26 Daji Landis , Nikolaj I. Schwartzbach

We study the complexity of problems related to subgame-perfect equilibria (SPEs) in infinite duration non zero-sum multiplayer games played on finite graphs with parity objectives. We present new complexity results that close gaps in the…

计算机科学与博弈论 · 计算机科学 2022-04-22 Léonard Brice , Marie van den Bogaard , Jean-François Raskin

We study nondeterministic strategies in parity games with the aim of computing a most permissive winning strategy. Following earlier work, we measure permissiveness in terms of the average number/weight of transitions blocked by the…

计算机科学中的逻辑 · 计算机科学 2013-01-14 Patricia Bouyer , Nicolas Markey , Jörg Olschewski , Michael Ummels

Prior beliefs are central to Bayesian accounts of cognition, but many of these accounts do not directly measure priors. More specifically, initial states of belief heavily influence how new information is assumed to be utilized when…

神经元与认知 · 定量生物学 2022-01-11 Peter A. V. DiBerardino , Alexandre L. S. Filipowicz , James Danckert , Britt Anderson

We introduce a "high probability" framework for repeated games with incomplete information. In our non-equilibrium setting, players aim to guarantee a certain payoff with high probability, rather than in expected value. We provide a high…

计算机科学与博弈论 · 计算机科学 2015-09-30 Payam Delgosha , Amin Gohari , Mohammad Akbarpour

When modifying existing policies in high-risk settings, it is often necessary to ensure with high certainty that the newly proposed policy improves upon a baseline, such as the status quo. In this work, we consider the problem of safe…

机器学习 · 计算机科学 2024-08-23 Brian M Cho , Ana-Roxana Pop , Kyra Gan , Sam Corbett-Davies , Israel Nir , Ariel Evnine , Nathan Kallus

We study the problem of Safe Policy Improvement (SPI) under constraints in the offline Reinforcement Learning (RL) setting. We consider the scenario where: (i) we have a dataset collected under a known baseline policy, (ii) multiple reward…

机器学习 · 计算机科学 2021-11-01 Harsh Satija , Philip S. Thomas , Joelle Pineau , Romain Laroche

We study stochastic two-player turn-based games in which the objective of one player is to ensure several infinite-horizon total reward objectives, while the other player attempts to spoil at least one of the objectives. The games have…

计算机科学与博弈论 · 计算机科学 2016-05-13 Romain Brenguier , Vojtěch Forejt

Secure equilibrium is a refinement of Nash equilibrium, which provides some security to the players against deviations when a player changes his strategy to another best response strategy. The concept of secure equilibrium is specifically…

计算机科学与博弈论 · 计算机科学 2014-05-08 Julie De Pril , János Flesch , Jeroen Kuipers , Gijs Schoenmakers , Koos Vrieze

While discounted payoff games and classic games that reduce to them, like parity and mean-payoff games, are symmetric, their solutions are not. We have taken a fresh view on the properties that optimal solutions need to have, and devised a…

数据结构与算法 · 计算机科学 2026-03-11 Daniele Dell'Erba , Arthur Dumas , Sven Schewe

There is a long history in game theory on the topic of Bayesian or "rational" learning, in which each player maintains beliefs over a set of alternative behaviours, or types, for the other players. This idea has gained increasing interest…

人工智能 · 计算机科学 2016-03-03 Stefano V. Albrecht , Jacob W. Crandall , Subramanian Ramamoorthy

Symmetry is inherent in the definition of most of the two-player zero-sum games, including parity, mean-payoff, and discounted-payoff games. It is therefore quite surprising that no symmetric analysis techniques for these games exist. We…

计算机科学与博弈论 · 计算机科学 2015-01-27 Sven Schewe , Ashutosh Trivedi , Thomas Varghese

We consider concurrent games played on graphs. At every round of a game, each player simultaneously and independently selects a move; the moves jointly determine the transition to a successor state. Two basic objectives are the safety…

计算机科学与博弈论 · 计算机科学 2012-07-03 Krishnendu Chatterjee , Luca de Alfaro , Thomas A. Henzinger

The strategy improvement algorithm for mean payoff games and parity games is a local improvement algorithm, just like the simplex algorithm for linear programs. Their similarity has turned out very useful: many lower bounds on running time…

计算机科学与博弈论 · 计算机科学 2025-09-22 Matthew Maat

We introduce the study of search games between a mobile Searcher and an immobile Hider in a new setting in which the Searcher has some potentially erroneous information, i.e., a prediction on the Hider's position. The objective is to…

计算机科学与博弈论 · 计算机科学 2024-09-05 Spyros Angelopoulos , Thomas Lidbetter , Konstantinos Panagiotou

We introduce the first complete formal solution to corrigibility in the off-switch game, with provable guarantees in multi-step, partially observed environments. Our framework consists of five *structurally separate* utility heads --…

人工智能 · 计算机科学 2025-11-20 Aran Nayebi