English
Related papers

Related papers: Core Safety Values for Provably Corrigible Agents

200 papers

*Automated circuit discovery* is a central tool in mechanistic interpretability for identifying the internal components of neural networks responsible for specific behaviors. While prior methods have made significant progress, they…

Machine Learning · Computer Science 2026-02-20 Itamar Hadad , Guy Katz , Shahaf Bassan

The holiday gift exchange game is a familiar social institution with nontrivial strategic structure. We provide a formal treatment of the game's mechanics, defining the state space, action sets, and the recursive structure of stealing…

Computer Science and Game Theory · Computer Science 2026-04-08 Daniel Quigley

There has been growing progress on theoretical analyses for provably efficient learning in MDPs with linear function approximation, but much of the existing work has made strong assumptions to enable exploration by conventional exploration…

Machine Learning · Computer Science 2020-10-23 Andrea Zanette , Alessandro Lazaric , Mykel J. Kochenderfer , Emma Brunskill

This work investigates the challenge of ensuring safety guarantees in the presence of uncontrollable agents, whose behaviors are stochastic and depend on both their own and the system's states. We present a neural model predictive control…

Systems and Control · Electrical Eng. & Systems 2026-04-21 Shuqi Wang , Mingyang Feng , Yu Chen , Yue Gao , Xiang Yin

Additively separable hedonic games and fractional hedonic games have received considerable attention. They are coalition forming games of selfish agents based on their mutual preferences. Most of the work in the literature characterizes the…

Artificial Intelligence · Computer Science 2017-06-29 Michele Flammini , Gianpiero Monaco , Qiang Zhang

Ensuring responsible use of artificial intelligence (AI) has become imperative as autonomous systems increasingly influence critical societal domains. However, the concept of trustworthy AI remains broad and multi-faceted. This thesis…

Artificial Intelligence · Computer Science 2025-10-28 Filip Cano

In this paper, we consider the setting of piecewise i.i.d. bandits under a safety constraint. In this piecewise i.i.d. setting, there exists a finite number of changepoints where the mean of some or all arms change simultaneously. We…

Machine Learning · Computer Science 2022-05-30 Subhojyoti Mukherjee

The framework of uncoupled online learning in multiplayer games has made significant progress in recent years. In particular, the development of time-varying games has considerably expanded its modeling capabilities. However, current regret…

Computer Science and Game Theory · Computer Science 2025-08-18 Aymeric Capitaine , Etienne Boursier , Eric Moulines , Michael I. Jordan , Alain Durmus

Motivated by the success of bounded model checking framework for finite state machines, Ouaknine and Worrell proposed a time-bounded theory of real-time verification by claiming that restriction to bounded-time recovers decidability for…

Logic in Computer Science · Computer Science 2014-08-18 Shankara Narayanan Krishna , Lakshmi Manasa , Ashutosh Trivedi

This paper presents a novel methodology to enforce motion safety guarantees even in the event of a sudden loss of control capabilities by any agent within a multi-agent system. This passive safety methodology permits the replacement of…

Optimization and Control · Mathematics 2023-05-29 Tommaso Guffanti , Simone D'Amico

In this paper we consider multi-agent coalitional games with uncertain value functions for which we establish distribution-free guarantees on the probability of allocation stability, i.e., agents do not have incentives to defect from the…

Optimization and Control · Mathematics 2022-06-24 George Pantazis , Filippo Fabiani , Filiberto Fele , Kostas Margellos

Embodied AI systems, comprising AI models and physical plants, are increasingly prevalent across various applications. Due to the rarity of system failures, ensuring their safety in complex operating environments remains a major challenge,…

In real-life scenarios, a Reinforcement Learning (RL) agent aiming to maximise their reward, must often also behave in a safe manner, including at training time. Thus, much attention in recent years has been given to Safe RL, where an agent…

Machine Learning · Statistics 2025-03-26 Edwin Hamel-De le Court , Francesco Belardinelli , Alexander W. Goodall

We consider the problem of adaptive control of a class of feedback linearizable plants with matched parametric uncertainties whose states are accessible, subject to state constraints, which often arise due to safety considerations. In this…

Systems and Control · Electrical Eng. & Systems 2026-01-13 Peter A. Fisher , Johannes Autenrieb , Anuradha M. Annaswamy

We consider the setting of stochastic multiagent systems modelled as stochastic multiplayer games and formulate an automated verification framework for quantifying and reasoning about agents' trust. To capture human trust, we work with a…

Logic in Computer Science · Computer Science 2019-05-17 Xiaowei Huang , Marta Kwiatkowska , Maciej Olejnik

Approachability has become a standard tool in analyzing earning algorithms in the adversarial online learning setup. We develop a variant of approachability for games where there is ambiguity in the obtained reward that belongs to a set,…

Statistics Theory · Mathematics 2012-02-17 Shie Mannor , Vianney Perchet , Gilles Stoltz

State of the art reinforcement learning methods sometimes encounter unsafe situations. Identifying when these situations occur is of interest both for post-hoc analysis and during deployment, where it might be advantageous to call out to a…

Machine Learning · Computer Science 2025-05-29 Alexander Grushin , Walt Woods , Alvaro Velasquez , Simon Khan

Obvious strategyproofness (OSP) is an appealing concept as it allows to maintain incentive compatibility even in the presence of agents that are not fully rational, e.g., those who struggle with contingent reasoning [Li, 2015]. However, it…

Computer Science and Game Theory · Computer Science 2017-02-21 Diodato Ferraioli , Carmine Ventre

There is a long history in game theory on the topic of Bayesian or "rational" learning, in which each player maintains beliefs over a set of alternative behaviours, or types, for the other players. This idea has gained increasing interest…

Artificial Intelligence · Computer Science 2016-03-03 Stefano V. Albrecht , Jacob W. Crandall , Subramanian Ramamoorthy

To achieve reliable, robust, and safe AI systems, it is vital to implement fallback strategies when AI predictions cannot be trusted. Certifiers for neural networks are a reliable way to check the robustness of these predictions. They…

Machine Learning · Computer Science 2023-10-04 Tobias Lorenz , Marta Kwiatkowska , Mario Fritz
‹ Prev 1 3 4 5 6 7 10 Next ›