English
Related papers

Related papers: Sustaining AI safety: Control-theoretic external i…

200 papers

What makes safety claims about general purpose AI systems such as large language models trustworthy? We show that rather than the capabilities of security tools such as alignment and red teaming procedures, it is security practices based on…

Cryptography and Security · Computer Science 2025-07-30 Petr Spelda , Vit Stritecky

Human oversight of AI is promoted as a safeguard against risks such as inaccurate outputs, system malfunctions, or violations of fundamental rights, and is mandated in regulation like the European AI Act. Yet debates on human oversight have…

Cryptography and Security · Computer Science 2026-03-06 Jonas C. Ditz , Veronika Lazar , Elmar Lichtmeß , Carola Plesch , Matthias Heck , Kevin Baum , Markus Langer

AI evaluations are an important component of the AI governance toolkit, underlying current approaches to safety cases for preventing catastrophic risks. Our paper examines what these evaluations can and cannot tell us. Evaluations can…

Computers and Society · Computer Science 2024-12-13 Peter Barnett , Lisa Thiergart

We investigate the important problem of certifying stability of reinforcement learning policies when interconnected with nonlinear dynamical systems. We show that by regulating the input-output gradients of policies, strong guarantees of…

Systems and Control · Computer Science 2018-10-30 Ming Jin , Javad Lavaei

Assuring safety for ``AI-based'' systems is one of the current challenges in safety engineering. For automated driving systems, in particular, further assurance challenges result from the open context that the systems need to operate in…

Systems and Control · Electrical Eng. & Systems 2025-07-29 Marcus Nolte , Nayel Fabian Salem , Olaf Franke , Jan Heckmann , Christoph Höhmann , Georg Stettinger , Markus Maurer

Artificial intelligence (AI) has been advancing at a fast pace and it is now poised for deployment in a wide range of applications, such as autonomous systems, medical diagnosis and natural language processing. Early adoption of AI…

Machine Learning · Computer Science 2023-09-21 Marta Kwiatkowska , Xiyue Zhang

The safety of mental health AI is often judged at the wrong temporal scale. Current evaluations typically score isolated responses, endpoint outcomes, or aggregate dialogue quality, while clinically consequential failures may arise from the…

Artificial Intelligence · Computer Science 2026-05-12 Srimonti Dutta , Ratna Kandala

Artificial intelligence systems are increasingly deployed in domains that shape human behaviour, institutional decision-making, and societal outcomes. Existing responsible AI and governance efforts provide important normative principles but…

Artificial Intelligence · Computer Science 2025-12-19 Otman A. Basir

Frontier AI Safety Policies concentrate on prevention: capability evaluations, deployment gates, and usage constraints, while neglecting the capacity to coordinate responses when prevention fails. We argue this coordination gap is…

Computers and Society · Computer Science 2026-05-21 Isaak Mengesha

Like any technology, AI systems come with inherent risks and potential benefits. It comes with potential disruption of established norms and methods of work, societal impacts and externalities. One may think of the adoption of technology as…

Computers and Society · Computer Science 2020-06-16 Mirka Snyder Caron , Abhishek Gupta

Different types of reasoning impose different structural demands on representational systems, yet no systematic account of these demands exists across psychology, AI, and philosophy of mind. I propose a framework identifying four structural…

Artificial Intelligence · Computer Science 2026-04-03 Yiling Wu

Artificial Intelligence (AI) methods are powerful tools for various domains, including critical fields such as avionics, where certification is required to achieve and maintain an acceptable level of safety. General solutions for…

We present a framework for characterizing neurosis in embodied AI: behaviors that are internally coherent yet misaligned with reality, arising from interactions among planning, uncertainty handling, and aversive memory. In a grid navigation…

Artificial Intelligence · Computer Science 2025-10-14 Daniel Howard

Advanced AI models hold the promise of tremendous benefits for humanity, but society needs to proactively manage the accompanying risks. In this paper, we focus on what we term "frontier AI" models: highly capable foundation models that…

Rapid progress in machine learning and artificial intelligence (AI) has brought increasing attention to the potential impacts of AI technologies on society. In this paper we discuss one such potential impact: the problem of accidents in…

Artificial Intelligence · Computer Science 2016-07-26 Dario Amodei , Chris Olah , Jacob Steinhardt , Paul Christiano , John Schulman , Dan Mané

The field of AI safety seeks to prevent or reduce the harms caused by AI systems. A simple and appealing account of what is distinctive of AI safety as a field holds that this feature is constitutive: a research project falls within the…

Computers and Society · Computer Science 2025-05-06 Jacqueline Harding , Cameron Domenico Kirk-Giannini

Critical infrastructure increasingly incorporates embodied AI for monitoring, predictive maintenance, and decision support. However, AI systems designed to handle statistically representable uncertainty struggle with cascading failures and…

Artificial Intelligence · Computer Science 2026-03-18 Puneet Sharma , Christer Henrik Pursiainen

Is there a way to design powerful AI systems based on machine learning methods that would satisfy probabilistic safety guarantees? With the long-term goal of obtaining a probabilistic guarantee that would apply in every context, we consider…

Artificial Intelligence · Computer Science 2025-06-17 Yoshua Bengio , Michael K. Cohen , Nikolay Malkin , Matt MacDermott , Damiano Fornasiere , Pietro Greiner , Younesse Kaddar

Public attention towards explainability of artificial intelligence (AI) systems has been rising in recent years to offer methodologies for human oversight. This has translated into the proliferation of research outputs, such as from…

Computers and Society · Computer Science 2023-04-25 Luca Nannini , Agathe Balayn , Adam Leon Smith

The accelerating displacement of human labor by artificial intelligence (AI) and robotic systems represents a structural transformation whose societal consequences extend far beyond conventional labor market analysis. This paper presents a…

Computers and Society · Computer Science 2026-04-02 Richard J. Mitchell
‹ Prev 1 4 5 6 7 8 10 Next ›