中文
相关论文

相关论文: Sustaining AI safety: Control-theoretic external i…

200 篇论文

What makes safety claims about general purpose AI systems such as large language models trustworthy? We show that rather than the capabilities of security tools such as alignment and red teaming procedures, it is security practices based on…

密码学与安全 · 计算机科学 2025-07-30 Petr Spelda , Vit Stritecky

Human oversight of AI is promoted as a safeguard against risks such as inaccurate outputs, system malfunctions, or violations of fundamental rights, and is mandated in regulation like the European AI Act. Yet debates on human oversight have…

密码学与安全 · 计算机科学 2026-03-06 Jonas C. Ditz , Veronika Lazar , Elmar Lichtmeß , Carola Plesch , Matthias Heck , Kevin Baum , Markus Langer

AI evaluations are an important component of the AI governance toolkit, underlying current approaches to safety cases for preventing catastrophic risks. Our paper examines what these evaluations can and cannot tell us. Evaluations can…

计算机与社会 · 计算机科学 2024-12-13 Peter Barnett , Lisa Thiergart

We investigate the important problem of certifying stability of reinforcement learning policies when interconnected with nonlinear dynamical systems. We show that by regulating the input-output gradients of policies, strong guarantees of…

系统与控制 · 计算机科学 2018-10-30 Ming Jin , Javad Lavaei

Assuring safety for ``AI-based'' systems is one of the current challenges in safety engineering. For automated driving systems, in particular, further assurance challenges result from the open context that the systems need to operate in…

系统与控制 · 电气工程与系统科学 2025-07-29 Marcus Nolte , Nayel Fabian Salem , Olaf Franke , Jan Heckmann , Christoph Höhmann , Georg Stettinger , Markus Maurer

Artificial intelligence (AI) has been advancing at a fast pace and it is now poised for deployment in a wide range of applications, such as autonomous systems, medical diagnosis and natural language processing. Early adoption of AI…

机器学习 · 计算机科学 2023-09-21 Marta Kwiatkowska , Xiyue Zhang

The safety of mental health AI is often judged at the wrong temporal scale. Current evaluations typically score isolated responses, endpoint outcomes, or aggregate dialogue quality, while clinically consequential failures may arise from the…

人工智能 · 计算机科学 2026-05-12 Srimonti Dutta , Ratna Kandala

Artificial intelligence systems are increasingly deployed in domains that shape human behaviour, institutional decision-making, and societal outcomes. Existing responsible AI and governance efforts provide important normative principles but…

人工智能 · 计算机科学 2025-12-19 Otman A. Basir

Frontier AI Safety Policies concentrate on prevention: capability evaluations, deployment gates, and usage constraints, while neglecting the capacity to coordinate responses when prevention fails. We argue this coordination gap is…

计算机与社会 · 计算机科学 2026-05-21 Isaak Mengesha

Like any technology, AI systems come with inherent risks and potential benefits. It comes with potential disruption of established norms and methods of work, societal impacts and externalities. One may think of the adoption of technology as…

计算机与社会 · 计算机科学 2020-06-16 Mirka Snyder Caron , Abhishek Gupta

Different types of reasoning impose different structural demands on representational systems, yet no systematic account of these demands exists across psychology, AI, and philosophy of mind. I propose a framework identifying four structural…

人工智能 · 计算机科学 2026-04-03 Yiling Wu

Artificial Intelligence (AI) methods are powerful tools for various domains, including critical fields such as avionics, where certification is required to achieve and maintain an acceptable level of safety. General solutions for…

We present a framework for characterizing neurosis in embodied AI: behaviors that are internally coherent yet misaligned with reality, arising from interactions among planning, uncertainty handling, and aversive memory. In a grid navigation…

人工智能 · 计算机科学 2025-10-14 Daniel Howard

Advanced AI models hold the promise of tremendous benefits for humanity, but society needs to proactively manage the accompanying risks. In this paper, we focus on what we term "frontier AI" models: highly capable foundation models that…

Rapid progress in machine learning and artificial intelligence (AI) has brought increasing attention to the potential impacts of AI technologies on society. In this paper we discuss one such potential impact: the problem of accidents in…

人工智能 · 计算机科学 2016-07-26 Dario Amodei , Chris Olah , Jacob Steinhardt , Paul Christiano , John Schulman , Dan Mané

The field of AI safety seeks to prevent or reduce the harms caused by AI systems. A simple and appealing account of what is distinctive of AI safety as a field holds that this feature is constitutive: a research project falls within the…

计算机与社会 · 计算机科学 2025-05-06 Jacqueline Harding , Cameron Domenico Kirk-Giannini

Critical infrastructure increasingly incorporates embodied AI for monitoring, predictive maintenance, and decision support. However, AI systems designed to handle statistically representable uncertainty struggle with cascading failures and…

人工智能 · 计算机科学 2026-03-18 Puneet Sharma , Christer Henrik Pursiainen

Is there a way to design powerful AI systems based on machine learning methods that would satisfy probabilistic safety guarantees? With the long-term goal of obtaining a probabilistic guarantee that would apply in every context, we consider…

Public attention towards explainability of artificial intelligence (AI) systems has been rising in recent years to offer methodologies for human oversight. This has translated into the proliferation of research outputs, such as from…

计算机与社会 · 计算机科学 2023-04-25 Luca Nannini , Agathe Balayn , Adam Leon Smith

The accelerating displacement of human labor by artificial intelligence (AI) and robotic systems represents a structural transformation whose societal consequences extend far beyond conventional labor market analysis. This paper presents a…

计算机与社会 · 计算机科学 2026-04-02 Richard J. Mitchell