中文
相关论文

相关论文: Safety is Non-Compositional: A Formal Framework fo…

200 篇论文

Compositionality is a critical aspect of scalable system design. Reinforcement learning (RL) has recently shown substantial success in task learning, but has only recently begun to truly leverage composition. In this paper, we focus on…

机器学习 · 计算机科学 2023-06-30 Kevin Leahy , Makai Mann , Zachary Serlin

Compositionality is believed to be fundamental to intelligence. In humans, it underlies the structure of thought, language, and higher-level reasoning. In AI, compositional representations can enable a powerful form of out-of-distribution…

计算与语言 · 计算机科学 2025-06-04 Eric Elmoznino , Thomas Jiralerspong , Yoshua Bengio , Guillaume Lajoie

This chapter argues that the reliability of agentic and generative AI is chiefly an architectural property. We define agentic systems as goal-directed, tool-using decision makers operating in closed loops, and show how reliability emerges…

人工智能 · 计算机科学 2025-12-11 Sławomir Nowaczyk

Safety cases, structured arguments that a system is acceptably safe, are becoming central to the governance of AI systems. Yet, traditional safety-case practices from aviation or nuclear engineering rely on well-specified system boundaries,…

软件工程 · 计算机科学 2026-03-09 Sung Une Lee , Liming Zhu , Md Shamsujjoha , Liming Dong , Qinghua Lu , Jieshan Chen , Lionel Briand

Generative artificial intelligence systems increasingly participate in research, law, education, media, and governance. Their fluent and adaptive outputs create an experience of collaboration. However, these systems do not bear…

人机交互 · 计算机科学 2026-02-24 Tatia Codreanu

Artificial intelligence (AI) is interacting with people at an unprecedented scale, offering new avenues for immense positive impact, but also raising widespread concerns around the potential for individual and societal harm. Today, the…

人工智能 · 计算机科学 2024-06-25 Andrea Bajcsy , Jaime F. Fisac

We examine two properties of AI systems: capability (what a system can do) and steerability (how reliably one can shift behavior toward intended outcomes). A central question is whether capability growth reduces steerability and risks…

计算与语言 · 计算机科学 2026-01-07 Jakub Hoscilowicz

Collaboration in multi-agent autonomous systems is critical to increase performance while ensuring safety. However, due to heterogeneity of their features in, e.g., perception qualities, some autonomous systems have to be considered more…

多智能体系统 · 计算机科学 2023-05-22 Selma Saidi

Existing accountability frameworks for AI systems, legal, ethical, and regulatory, rest on a shared assumption: for any consequential outcome, at least one identifiable person had enough involvement and foresight to bear meaningful…

人工智能 · 计算机科学 2026-04-10 Haileleol Tibebu

Current agentic AI architectures are fundamentally incompatible with the security and epistemological requirements of high-stakes scientific workflows. The problem is not inadequate alignment or insufficient guardrails, it is architectural:…

密码学与安全 · 计算机科学 2026-02-11 Manish Bhattarai , Minh Vu

The design of embedded safety-critical systems such as those used in next-generation automotive and autonomous platforms, is increasingly challenged by escalating system complexity, hardware-software heterogeneity, and the integration of…

It has been increasingly recognized that effective human-AI co-creation requires more than prompts and results, but an environment with empowering structures that facilitate exploration, planning, iteration, as well as control and…

人机交互 · 计算机科学 2025-03-07 Yining Cao , Yiyi Huang , Anh Truong , Hijung Valentina Shin , Haijun Xia

Recent advances in AI research make it increasingly plausible that artificial agents with consequential real-world impact will soon operate beyond tightly controlled environments. Ensuring that these agents are not only safe but that they…

计算机与社会 · 计算机科学 2025-06-10 Kevin Baum

Insofar as consciousness has a functional role in facilitating learning and behavioral control, the builders of autonomous AI systems are likely to attempt to incorporate it into their designs. The extensive literature on the ethics of AI…

计算机与社会 · 计算机科学 2020-02-14 Aman Agarwal , Shimon Edelman

This paper advances a methodological proposal for safety research in agentic AI. As systems acquire planning, memory, tool use, persistent identity, and sustained interaction, safety can no longer be analysed primarily at the level of the…

计算机与社会 · 计算机科学 2026-04-17 Federico Pierucci , Matteo Prandi , Marcantonio Bracale Syrnikov , Marcello Galisai , Piercosma Bisconti

Safety is an important element of dependability. It is defined as the absence of accidents. Most accidents involving software-intensive systems have been system accidents, which are caused by unsafe inter-system or inter-component…

软件工程 · 计算机科学 2016-11-17 Zhe Chen , Gilles Motet

With edge-AI finding an increasing number of real-world applications, especially in industry, the question of functionally safe applications using AI has begun to be asked. In this body of work, we explore the issue of achieving dependable…

机器学习 · 计算机科学 2021-08-06 Hans Dermot Doran , Gianluca Ielpo , David Ganz , Michael Zapke

Generative AI systems are increasingly assisting and acting on behalf of end users in practical settings, from digital shopping assistants to next-generation autonomous cars. In this context, safety is no longer about blocking harmful…

人工智能 · 计算机科学 2026-05-20 Ravi Pandya , Madison Bland , Duy P. Nguyen , Changliu Liu , Jaime Fernández Fisac , Andrea Bajcsy

AI advancements have been significantly driven by a combination of foundation models and curiosity-driven learning aimed at increasing capability and adaptability. Within this landscape, open-endedness, where AI agents autonomously and…

人工智能 · 计算机科学 2026-05-06 Ivaxi Sheth , Jan Wehner , Sahar Abdelnabi , Ruta Binkyte , Mario Fritz

We introduce the first complete formal solution to corrigibility in the off-switch game, with provable guarantees in multi-step, partially observed environments. Our framework consists of five *structurally separate* utility heads --…

人工智能 · 计算机科学 2025-11-20 Aran Nayebi