中文
相关论文

相关论文: Sustaining AI safety: Control-theoretic external i…

200 篇论文

The exposure of security vulnerabilities in safety-aligned language models, e.g., susceptibility to adversarial attacks, has shed light on the intricate interplay between AI safety and AI security. Although the two disciplines now come…

This chapter formulates seven lessons for preventing harm in artificial intelligence (AI) systems based on insights from the field of system safety for software-based automation in safety-critical domains. New applications of AI across…

系统与控制 · 电气工程与系统科学 2022-02-21 Roel I. J. Dobbe

Ensuring responsible use of artificial intelligence (AI) has become imperative as autonomous systems increasingly influence critical societal domains. However, the concept of trustworthy AI remains broad and multi-faceted. This thesis…

人工智能 · 计算机科学 2025-10-28 Filip Cano

It's widely expected that humanity will someday create AI systems vastly more intelligent than us, leading to the unsolved alignment problem of "how to control superintelligence." However, this commonly expressed problem is not only…

人工智能 · 计算机科学 2024-12-02 James M. Mazzu

As Artificial Intelligence (AI) systems increasingly assume consequential decision-making roles, a widening gap has emerged between technical capabilities and institutional accountability. Ethical guidance alone is insufficient to counter…

人工智能 · 计算机科学 2025-12-12 Byeong Ho Kang , Wenli Yang , Muhammad Bilal Amin

This vision paper presents initial research on assessing the robustness and reliability of AI-enabled systems, and key factors in ensuring their safety and effectiveness in practical applications, including a focus on accountability. By…

软件工程 · 计算机科学 2025-06-23 Filippo Scaramuzza , Damian A. Tamburri , Willem-Jan van den Heuvel

How can we ensure that AI systems are aligned with human values and remain safe? We can study this problem through the frameworks of the AI assistance and the AI shutdown games. The AI assistance problem concerns designing an AI agent that…

人工智能 · 计算机科学 2025-12-30 Alessio Benavoli , Alessandro Facchini , Marco Zaffalon

This paper explores the role and challenges of Artificial Intelligence (AI) algorithms, specifically AI-based software elements, in autonomous driving systems. These AI systems are fundamental in executing real-time critical functions in…

人工智能 · 计算机科学 2024-03-01 Mandar Pitale , Alireza Abbaspour , Devesh Upadhyay

We draw on our experience working on system and software assurance and evaluation for systems important to society to summarise how safety engineering is performed in traditional critical systems, such as aircraft flight control. We analyse…

计算机与社会 · 计算机科学 2025-02-07 Robin Bloomfield , John Rushby

As AI systems advance, AI evaluations are becoming an important pillar of regulations for ensuring safety. We argue that such regulation should require developers to explicitly identify and justify key underlying assumptions about…

人工智能 · 计算机科学 2024-11-21 Peter Barnett , Lisa Thiergart

Open-source software (OSS) is foundational to modern digital infrastructure, yet this context for group work continues to struggle to ensure sufficient contributions in many critical cases. This literature review explores how artificial…

软件工程 · 计算机科学 2026-02-10 S M Rakib UI Karim , Wenyi Lu , Sean Goggins

As large language models are increasingly deployed as interacting agents in high-stakes decisions, the AI safety community assumes that safety properties of individual models will compose into safe multi-agent behavior. This position paper…

人工智能 · 计算机科学 2026-05-05 Tanav Singh Bajaj , Nikhil Singh , Karan Anand , Eishkaran Singh

AI safety has emerged as a critical priority as these systems are increasingly deployed in real-world applications. We propose the first domain-agnostic AI safety ensuring framework that achieves strong safety guarantees while preserving…

人工智能 · 计算机科学 2025-10-07 Beomjun Kim , Kangyeon Kim , Sunwoo Kim , Yeonsang Shin , Heejin Ahn

The increasing reliance on AI-driven 5G/6G network infrastructures for mission-critical services highlights the need for reliability and resilience against sophisticated cyber-physical threats. These networks are highly exposed to novel…

The study of complex adaptive systems, pioneered in physics, biology, and the social sciences, offers important lessons for AI governance. Contemporary AI systems and the environments in which they operate exhibit many of the properties…

计算机与社会 · 计算机科学 2025-03-04 Noam Kolt , Michal Shur-Ofry , Reuven Cohen

Robustness is key to engineering, automation, and science as a whole. However, the property of robustness is often underpinned by costly requirements such as over-provisioning, known uncertainty and predictive models, and known adversaries.…

机器人学 · 计算机科学 2021-09-28 Amanda Prorok , Matthew Malencia , Luca Carlone , Gaurav S. Sukhatme , Brian M. Sadler , Vijay Kumar

Frontier AI developers operate at the intersection of rapid technical progress, extreme risk exposure, and growing regulatory scrutiny. While a range of external evaluations and safety frameworks have emerged, comparatively little attention…

计算机与社会 · 计算机科学 2025-12-19 Francesca Gomez , Adam Buick , Leah Ferentinos , Haelee Kim , Elley Lee

Explainable Artificial Intelligence (XAI) is increasingly rec ognized as essential for deploying machine learning systems in safety critical environments. In Automatic Target Recognition (ATR), where models operate on image, video, radar,…

人工智能 · 计算机科学 2026-05-08 Vanessa Buhrmester , David Muench , Dimitri Bulatov , Michael Arens

Risk-based AI regulation has become the dominant paradigm in AI governance, promising proportional controls aligned with anticipated harms. This paper argues that such frameworks often fail for structural reasons: they implicitly assume…

计算机与社会 · 计算机科学 2025-12-16 Hugo Roger Paz

Understanding AI systems' inner workings is critical for ensuring value alignment and safety. This review explores mechanistic interpretability: reverse engineering the computational mechanisms and representations learned by neural networks…

人工智能 · 计算机科学 2024-08-27 Leonard Bereska , Efstratios Gavves