English
Related papers

Related papers: Promises Made, Promises Kept: Safe Pareto Improvem…

200 papers

Agents in mixed-motive coordination problems such as Chicken may fail to coordinate on a Pareto-efficient outcome. Safe Pareto improvements (SPIs) were originally proposed to mitigate miscoordination in cases where players lack…

Computer Science and Game Theory · Computer Science 2025-02-24 Anthony DiGiovanni , Jesse Clifton , Nicolas Macé

In an offline reinforcement learning setting, the safe policy improvement (SPI) problem aims to improve the performance of a behavior policy according to which sample data has been generated. State-of-the-art approaches to SPI require a…

Machine Learning · Computer Science 2023-05-16 Patrick Wienhöft , Marnix Suilen , Thiago D. Simão , Clemens Dubslaff , Christel Baier , Nils Jansen

We consider a setting in which a principal gets to choose which game from some given set is played by a group of agents. The principal would like to choose a game that favors one of the players, the social preferences of the players, or the…

Computer Science and Game Theory · Computer Science 2025-11-27 Caspar Oesterheld , Vincent Conitzer

Safe policy improvement (SPI) is an offline reinforcement learning problem in which a new policy that reliably outperforms the behavior policy with high confidence needs to be computed using only a dataset and the behavior policy. Markov…

Artificial Intelligence · Computer Science 2025-08-20 Kasper Engelen , Guillermo A. Pérez , Marnix Suilen

Safe Policy Improvement (SPI) aims at provable guarantees that a learned policy is at least approximately as good as a given baseline policy. Building on SPI with Soft Baseline Bootstrapping (Soft-SPIBB) by Nadjahi et al., we identify…

Machine Learning · Computer Science 2022-08-02 Philipp Scholl , Felix Dietrich , Clemens Otte , Steffen Udluft

Safe Policy Improvement (SPI) is an important technique for offline reinforcement learning in safety critical applications as it improves the behavior policy with a high probability. We classify various SPI approaches from the literature…

Machine Learning · Computer Science 2022-08-02 Philipp Scholl , Felix Dietrich , Clemens Otte , Steffen Udluft

Safe policy improvement (SPI) offers theoretical control over policy updates, yet existing guarantees largely concern offline, tabular reinforcement learning (RL). We study SPI in general online settings, when combined with world model and…

Machine Learning · Computer Science 2026-01-29 Florent Delgrange , Raphael Avalos , Willem Röpke

We study discounted infinitely repeated games in which players agree on a cooperative mixed action profile but, at each step, observe only the realized pure actions. This form of imperfect monitoring breaks classical trigger strategies,…

Applications · Statistics 2026-03-09 Aymeric Capitaine , Antoine Scheid , Etienne Boursier , Alain Durmus , Michael I. Jordan

The combination of the Bayesian game and learning has a rich history, with the idea of controlling a single agent in a system composed of multiple agents with unknown behaviors given a set of types, each specifying a possible behavior for…

Machine Learning · Computer Science 2024-11-21 Tongxin Li , Tinashe Handina , Shaolei Ren , Adam Wierman

We solve the problem of automatically computing a new class of environment assumptions in two-player turn-based finite graph games which characterize an ``adequate cooperation'' needed from the environment to allow the system player to win.…

Computer Science and Game Theory · Computer Science 2024-01-23 Ashwani Anand , Kaushik Mallik , Satya Prakash Nayak , Anne-Kathrin Schmuck

Simple stochastic games are turn-based 2.5-player games with a reachability objective. The basic question asks whether one player can ensure reaching a given target with at least a given probability. A natural extension is games with a…

Computer Science and Game Theory · Computer Science 2021-02-02 Pranav Ashok , Krishnendu Chatterjee , Jan Kretinsky , Maximilian Weininger , Tobias Winkler

In offline reinforcement learning (RL), we learn policies from fixed datasets without environment interaction. The major challenges are to provide guarantees on the (1) performance and (2) safety of the resulting policy. A technique called…

Machine Learning · Computer Science 2026-05-12 Maris F. L. Galesloot , Thomas Rhemrev , Nils Jansen

This paper studies two-player zero-sum repeated Bayesian games in which every player has a private type that is unknown to the other player, and the initial probability of the type of every player is publicly known. The types of players are…

Computer Science and Game Theory · Computer Science 2017-11-08 Lichun Li , Cedric Langbort , Jeff Shamma

In imperfect information games (e.g. Bridge, Skat, Poker), one of the fundamental considerations is to infer the missing information while at the same time avoiding the disclosure of private information. Disregarding the issue of protecting…

Artificial Intelligence · Computer Science 2024-05-24 Jérôme Arjonilla , Abdallah Saffidine , Tristan Cazenave

Preferences play a key role in determining what goals/constraints to satisfy when not all constraints can be satisfied simultaneously. In this paper, we study how to synthesize preference satisfying plans in stochastic systems, modeled as…

Artificial Intelligence · Computer Science 2022-10-06 Abhishek N. Kulkarni , Jie Fu

The Type Indeterminacy model is a theoretical framework that formalizes the constructive preference perspective suggested by Kahneman and Tversky. In this paper we explore an extention of the TI-model from simple to strategic…

Physics and Society · Physics 2009-02-11 Jerome Busemeyer , A. Lambert-Mogiliansky

We address two-player general-sum stochastic Stackelberg games (SSGs), where the leader's policy is optimized considering the best-response follower whose policy is optimal for its reward under the leader. Existing policy gradient and value…

Computer Science and Game Theory · Computer Science 2026-03-17 Mikoto Kudo , Youhei Akimoto

We study safe policy improvement (SPI) for partially observable Markov decision processes (POMDPs). SPI is an offline reinforcement learning (RL) problem that assumes access to (1) historical data about an environment, and (2) the so-called…

Artificial Intelligence · Computer Science 2023-01-13 Thiago D. Simão , Marnix Suilen , Nils Jansen

Human expectations arise from their understanding of others and the world. In the context of human-AI interaction, this understanding may not align with reality, leading to the AI agent failing to meet expectations and compromising team…

Robotics · Computer Science 2024-04-01 Akkamahadevi Hanni , Andrew Boateng , Yu Zhang

Previous work has shown the unreliability of existing algorithms in the batch Reinforcement Learning setting, and proposed the theoretically-grounded Safe Policy Improvement with Baseline Bootstrapping (SPIBB) fix: reproduce the baseline…

Machine Learning · Computer Science 2021-01-01 Thiago D. Simão , Romain Laroche , Rémi Tachet des Combes
‹ Prev 1 2 3 10 Next ›