中文
相关论文

相关论文: Arguments about Highly Reliable Agent Designs as a…

200 篇论文

This article critically examines the recent hype around AI safety. We first start with noting the nature of the AI safety hype as being dominated by governments and corporations, and contrast it with other avenues within AI research on…

人工智能 · 计算机科学 2024-03-28 Deepak P

The Robust Artificial Intelligence System Assurance (RAISA) workshop will focus on research, development and application of robust artificial intelligence (AI) and machine learning (ML) systems. Rather than studying robustness with respect…

人工智能 · 计算机科学 2022-02-11 Olivia Brown , Brad Dillman

This paper examines the responsible integration of artificial intelligence (AI) in human services organizations (HSOs), proposing a nuanced framework for evaluating AI applications across multiple dimensions of risk. The authors argue that…

计算机与社会 · 计算机科学 2025-01-22 Brian E. Perron , Lauri Goldkind , Zia Qi , Bryan G. Victor

We present a suite of reinforcement learning environments illustrating various safety properties of intelligent agents. These problems include safe interruptibility, avoiding side effects, absent supervisor, reward gaming, safe exploration,…

AI agents, predominantly powered by large language models (LLMs), are vulnerable to indirect prompt injection, in which malicious instructions embedded in untrusted data can trigger dangerous agent actions. This position paper discusses our…

密码学与安全 · 计算机科学 2026-04-01 Chong Xiang , Drew Zagieboylo , Shaona Ghosh , Sanjay Kariyappa , Kai Greshake , Hanshen Xiao , Chaowei Xiao , G. Edward Suh

Ensuring that AI systems reliably and robustly avoid harmful or dangerous behaviours is a crucial challenge, especially for AI systems with a high degree of autonomy and general intelligence, or systems used in safety-critical contexts. In…

Existing evaluations of AI misuse safeguards provide a patchwork of evidence that is often difficult to connect to real-world decisions. To bridge this gap, we describe an end-to-end argument (a "safety case") that misuse safeguards reduce…

机器学习 · 计算机科学 2025-05-26 Joshua Clymer , Jonah Weinbaum , Robert Kirk , Kimberly Mai , Selena Zhang , Xander Davies

As artificial intelligence (AI) becomes increasingly embedded in digital, social, and institutional infrastructures, and AI and platforms are merged into hybrid structures, systemic risk has emerged as a critical but undertheorized…

计算机与社会 · 计算机科学 2026-05-26 Philipp Hacker , Lilian Edwards , Atoosa Kasirzadeh

Is there a way to design powerful AI systems based on machine learning methods that would satisfy probabilistic safety guarantees? With the long-term goal of obtaining a probabilistic guarantee that would apply in every context, we consider…

The use of Artificial Intelligence (AI) in high-risk, decision-making scenarios presents technical, safety, and normative challenges; problems that may only be ameliorated by human oversight. However, notions of human oversight lack a…

Artificial Intelligence (AI) is an effective science which employs strong enough approaches, methods, and techniques to solve unsolvable real world based problems. Because of its unstoppable rise towards the future, there are also some…

人工智能 · 计算机科学 2017-06-12 Alice Pavaloiu , Utku Kose

Safety cases - clear, assessable arguments for the safety of a system in a given context - are a widely-used technique across various industries for showing a decision-maker (e.g. boards, customers, third parties) that a system is safe. In…

计算机与社会 · 计算机科学 2025-03-10 Benjamin Hilton , Marie Davidsen Buhl , Tomek Korbak , Geoffrey Irving

Explanations for artificial intelligence (AI) systems are intended to support the people who are impacted by AI systems in high-stakes decision-making environments, such as doctors, patients, teachers, students, housing applicants, and many…

人机交互 · 计算机科学 2025-04-16 Gennie Mansi , Naveena Karusala , Mark Riedl

Rapid progress in machine learning and artificial intelligence (AI) has brought increasing attention to the potential impacts of AI technologies on society. In this paper we discuss one such potential impact: the problem of accidents in…

人工智能 · 计算机科学 2016-07-26 Dario Amodei , Chris Olah , Jacob Steinhardt , Paul Christiano , John Schulman , Dan Mané

As artificial intelligence scales, the concepts of alignment, agency, and autonomy have become central to AI safety, governance, and control. However, even in human contexts, these terms lack universal definitions, varying across…

计算机与社会 · 计算机科学 2025-03-11 Krti Tallam

Automated driving (AD) is promising, but the transition to fully autonomous driving is, among other things, subject to the real, ever-changing open world and the resulting challenges. However, research in the field of AD demonstrates the…

机器人学 · 计算机科学 2026-02-02 Lars Ullrich , Michael Buchholz , Klaus Dietmayer , Knut Graichen

While artificial intelligence (AI) is advancing rapidly and mastering increasingly complex problems with astonishing performance, the safety assurance of such systems is a major concern. Particularly in the context of safety-critical,…

人工智能 · 计算机科学 2025-07-01 Lars Ullrich , Walter Zimmer , Ross Greer , Knut Graichen , Alois C. Knoll , Mohan Trivedi

To facilitate the widespread acceptance of AI systems guiding decision-making in real-world applications, it is key that solutions comprise trustworthy, integrated human-AI systems. Not only in safety-critical applications such as…

人工智能 · 计算机科学 2020-01-16 Florian Buettner , John Piorkowski , Ian McCulloh , Ulli Waltinger

Artificial Knowledge (AK) systems are transforming decision-making across critical domains such as healthcare, finance, and criminal justice. However, their growing opacity presents governance challenges that current regulatory approaches,…

计算机与社会 · 计算机科学 2025-05-29 Dalit Ken-Dror Feldman , Daniel Benoliel

Artificial Intelligence (AI) started out with an ambition to reproduce the human mind, but, as the sheer scale of that ambition became manifest, it quickly retreated into either studying specialized intelligent behaviours, or proposing…

人工智能 · 计算机科学 2021-06-17 Alexander Boer , Giovanni Sileno