English
Related papers

Related papers: Core Safety Values for Provably Corrigible Agents

200 papers

Deep learning is built on the foundational guarantee that gradient descent on an objective function converges to local minima. Unfortunately, this guarantee fails in settings, such as generative adversarial nets, that exhibit multiple…

Machine Learning · Computer Science 2019-05-14 Alistair Letcher , David Balduzzi , Sebastien Racaniere , James Martens , Jakob Foerster , Karl Tuyls , Thore Graepel

As artificial intelligence (AI) / machine learning (ML) gain widespread adoption, practitioners are increasingly seeking means to quantify and control the risk these systems incur. This challenge is especially salient when such systems have…

Machine Learning · Computer Science 2024-06-06 Drew Prinster , Samuel Stanton , Anqi Liu , Suchi Saria

The SmoothLLM defense provides a certification guarantee against jailbreaking attacks, but it relies on a strict "k-unstable" assumption that rarely holds in practice. This strong assumption can limit the trustworthiness of the provided…

Machine Learning · Computer Science 2026-03-10 Adarsh Kumarappan , Ayushi Mehrotra

This paper contains the first formal proof that safety is non-compositional in the presence of conjunctive capability dependencies: two agents each individually inca- pable of reaching any forbidden capability can, when combined,…

Artificial Intelligence · Computer Science 2026-03-20 Cosimo Spera

Safety guarantees are a prerequisite to the deployment of reinforcement learning (RL) agents in safety-critical tasks. Often, deployment environments exhibit non-stationary dynamics or are subject to changing performance goals, requiring…

Machine Learning · Computer Science 2026-04-13 Maksim Anisimov , Francesco Belardinelli , Matthew Wicker

Reinforcement learning (RL) is increasingly used to personalize instruction in intelligent tutoring systems, yet the field lacks a formal framework for defining and evaluating pedagogical safety. We introduce a four-layer model of…

Artificial Intelligence · Computer Science 2026-04-07 Oluseyi Olukola , Nick Rahimi

As LM agents gain the ability to execute actions on real computer systems, we need ways to not only prevent harmful actions at scale but also effectively remediate harm when prevention fails. We formalize a solution to this neglected…

Artificial Intelligence · Computer Science 2026-05-29 Christy Li , Sky CH-Wang , Andi Peng , Andreea Bobu

Partially observable Markov decision processes (POMDPs) form a prominent model for uncertainty in sequential decision making. We are interested in constructing algorithms with theoretical guarantees to determine whether the agent has a…

Artificial Intelligence · Computer Science 2024-12-17 Marius Belly , Nathanaël Fijalkow , Hugo Gimbert , Florian Horn , Guillermo A. Pérez , Pierre Vandenhove

Reinforcement learning (RL) systems typically optimize scalar reward functions that assume precise and reliable evaluation of outcomes. However, real-world objectives--especially those derived from human preferences--are often uncertain,…

Machine Learning · Computer Science 2026-04-30 Disha Singha

We study a sequential mechanism design problem in which a principal seeks to elicit truthful reports from multiple rational agents while starting with no prior knowledge of agents' beliefs. We introduce Distributionally Robust Adaptive…

Computer Science and Game Theory · Computer Science 2026-04-22 Qiushi Han , David Simchi-Levi , Renfei Tan , Zishuo Zhao

While single-agent policy optimization in a fixed environment has attracted a lot of research attention recently in the reinforcement learning community, much less is known theoretically when there are multiple agents playing in a…

Machine Learning · Computer Science 2022-07-27 Shuang Qiu , Xiaohan Wei , Jieping Ye , Zhaoran Wang , Zhuoran Yang

The deployment of autonomous AI agents in sensitive domains, such as healthcare, introduces critical risks to safety, security, and privacy. These agents may deviate from user objectives, violate data handling policies, or be compromised by…

Software Engineering · Computer Science 2025-10-08 Lesly Miculicich , Mihir Parmar , Hamid Palangi , Krishnamurthy Dj Dvijotham , Mirko Montanari , Tomas Pfister , Long T. Le

Recent AI systems compress the distance between capability growth and capability deployment. Earlier high-risk technologies were slowed by capital intensity, physical bottlenecks, organizational inertia, and specialized supply chains. By…

Artificial Intelligence · Computer Science 2026-05-05 Wesley Shu , Peng Wei

AI safety has emerged as a critical priority as these systems are increasingly deployed in real-world applications. We propose the first domain-agnostic AI safety ensuring framework that achieves strong safety guarantees while preserving…

Artificial Intelligence · Computer Science 2025-10-07 Beomjun Kim , Kangyeon Kim , Sunwoo Kim , Yeonsang Shin , Heejin Ahn

Safety assurance is a fundamental requirement for deploying learning-enabled autonomous systems. Hamilton-Jacobi (HJ) reachability analysis is a fundamental method for formally verifying safety and generating safe controllers. However,…

Machine Learning · Computer Science 2025-11-21 Ihab Tabbara , Yuxuan Yang , Hussein Sibai

Can generative agents be trusted in multimodal environments? Despite advances in large language and vision-language models that enable agents to act autonomously and pursue goals in rich settings, their ability to reason about safety,…

Artificial Intelligence · Computer Science 2025-10-10 Alhim Vera , Karen Sanchez , Carlos Hinojosa , Haidar Bin Hamid , Donghoon Kim , Bernard Ghanem

First-order (FO) transition systems have recently attracted attention for the verification of parametric systems such as network protocols, software-defined networks or multi-agent workflows like conference management systems. Desirable…

Logic in Computer Science · Computer Science 2019-11-15 Helmut Seidl , Christian Müller , Bernd Finkbeiner

Ttraditional safety engineering is coming to a turning point moving from deterministic, non-evolving systems operating in well-defined contexts to increasingly autonomous and learning-enabled AI systems which are acting in largely…

Artificial Intelligence · Computer Science 2022-05-13 Harald Rueß , Simon Burton

We present a complete theoretical characterization of Latent Posterior Factors (LPF), a principled framework for aggregating multiple heterogeneous evidence items in probabilistic prediction tasks. Multi-evidence reasoning arises…

Artificial Intelligence · Computer Science 2026-03-19 Aliyu Agboola Alege

Adjustable autonomy refers to entities dynamically varying their own autonomy, transferring decision-making control to other entities (typically agents transferring control to human users) in key situations. Determining whether and when…

Artificial Intelligence · Computer Science 2011-06-24 D. V. Pynadath , P. Scerri , M. Tambe
‹ Prev 1 8 9 10 Next ›