English
Related papers

Related papers: Can a Bayesian Oracle Prevent Harm from an Agent?

200 papers

Classical reinforcement learning assumes the agent interacts with a fixed environment whose behavior does not depend on the agent's policy. This assumption breaks down in non-realizable settings where other actors might anticipate the…

Human oversight of AI is promoted as a safeguard against risks such as inaccurate outputs, system malfunctions, or violations of fundamental rights, and is mandated in regulation like the European AI Act. Yet debates on human oversight have…

Cryptography and Security · Computer Science 2026-03-06 Jonas C. Ditz , Veronika Lazar , Elmar Lichtmeß , Carola Plesch , Matthias Heck , Kevin Baum , Markus Langer

Leveraging recent developments in black-box risk-aware verification, we provide three algorithms that generate probabilistic guarantees on (1) optimality of solutions, (2) recursive feasibility, and (3) maximum controller runtimes for…

Optimization and Control · Mathematics 2023-03-14 Prithvi Akella , Wyatt Ubellacker , Aaron D. Ames

The capabilities of artificial intelligence systems have been advancing to a great extent, but these systems still struggle with failure modes, vulnerabilities, and biases. In this paper, we study the current state of the field, and present…

Cryptography and Security · Computer Science 2025-06-12 Xingli Fang , Jianwei Li , Varun Mulchandani , Jung-Eun Kim

Artificial intelligence commonly refers to the science and engineering of artificial systems that can carry out tasks generally associated with requiring aspects of human intelligence, such as playing games, translating languages, and…

Artificial Intelligence · Computer Science 2025-02-11 Andreas Krause , Jonas Hübotter

In this work, we present and analyze reported failures of artificially intelligent systems and extrapolate our analysis to future AIs. We suggest that both the frequency and the seriousness of future AI failures will steadily increase. AI…

Artificial Intelligence · Computer Science 2016-10-26 Roman V. Yampolskiy , M. S. Spellchecker

Our intention is to provide a definitive reference on what it would take to safely make use of generative/predictive models in the absence of a solution to the Eliciting Latent Knowledge problem. Furthermore, we believe that large language…

Artificial Intelligence · Computer Science 2023-02-07 Evan Hubinger , Adam Jermyn , Johannes Treutlein , Rubi Hudson , Kate Woolverton

As machine learning systems become more powerful they also become increasingly unpredictable and opaque. Yet, finding human-understandable explanations of how they work is essential for their safe deployment. This technical report…

Shielding is a prominent model-based technique to ensure safety of autonomous agents. Classical shielding aims to ensure that nothing bad ever happens and comes with strong guarantees about safety and maximal permissiveness. However,…

Logic in Computer Science · Computer Science 2026-05-14 Linus Heck , Filip Macák , Roman Andriushchenko , Milan Češka , Sebastian Junges

This chapter formulates seven lessons for preventing harm in artificial intelligence (AI) systems based on insights from the field of system safety for software-based automation in safety-critical domains. New applications of AI across…

Systems and Control · Electrical Eng. & Systems 2022-02-21 Roel I. J. Dobbe

In this position paper, we address the persistent gap between rapidly growing AI capabilities and lagging safety progress. Existing paradigms divide into ``Make AI Safe'', which applies post-hoc alignment and guardrails but remains brittle…

Machine Learning · Computer Science 2025-09-09 Youbang Sun , Xiang Wang , Jie Fu , Chaochao Lu , Bowen Zhou

When learning policies for robotic systems from data, safety is a major concern, as violation of safety constraints may cause hardware damage. SafeOpt is an efficient Bayesian optimization (BO) algorithm that can learn policies while…

Robotics · Computer Science 2021-05-28 Dominik Baumann , Alonso Marco , Matteo Turchetta , Sebastian Trimpe

We draw on our experience working on system and software assurance and evaluation for systems important to society to summarise how safety engineering is performed in traditional critical systems, such as aircraft flight control. We analyse…

Computers and Society · Computer Science 2025-02-07 Robin Bloomfield , John Rushby

In the future, powerful AI systems may be deployed in high-stakes settings, where a single failure could be catastrophic. One technique for improving AI safety in high-stakes settings is adversarial training, which uses an adversary to…

As artificial intelligence systems grow more capable and autonomous, frontier AI development poses potential systemic risks that could affect society at a massive scale. Current practices at many AI labs developing these systems lack…

Computers and Society · Computer Science 2025-06-03 Aidan Kierans , Kaley Rittichier , Utku Sonsayar , Avijit Ghosh

Artificial Intelligence (AI) is rapidly being integrated into critical systems across various domains, from healthcare to autonomous vehicles. While its integration brings immense benefits, it also introduces significant risks, including…

Computers and Society · Computer Science 2025-06-25 Zhiqiang Lin , Huan Sun , Ness Shroff

We show that when a third party, the adversary, steps into the two-party setting (agent and operator) of safely interruptible reinforcement learning, a trade-off has to be made between the probability of following the optimal policy in the…

Machine Learning · Computer Science 2018-05-30 Henrik Aslund , El Mahdi El Mhamdi , Rachid Guerraoui , Alexandre Maurer

Large language model-based agents are rapidly evolving from simple conversational assistants into autonomous systems capable of performing complex, professional-level tasks in various domains. While these advancements promise significant…

This paper considers the problem of learning safe policies in the context of reinforcement learning (RL). In particular, we consider the notion of probabilistic safety. This is, we aim to design policies that maintain the state of the…

Machine Learning · Computer Science 2023-04-20 Weiqin Chen , Dharmashankar Subramanian , Santiago Paternain

AI safety has emerged as a critical priority as these systems are increasingly deployed in real-world applications. We propose the first domain-agnostic AI safety ensuring framework that achieves strong safety guarantees while preserving…

Artificial Intelligence · Computer Science 2025-10-07 Beomjun Kim , Kangyeon Kim , Sunwoo Kim , Yeonsang Shin , Heejin Ahn