中文
相关论文

相关论文: An alignment safety case sketch based on debate

200 篇论文

Background: Value alignment in computer science research is often used to refer to the process of aligning artificial intelligence with humans, but the way the phrase is used often lacks precision. Objectives: In this paper, we conduct a…

计算机与社会 · 计算机科学 2026-03-27 Jack McKinlay , Marina De Vos , Janina A. Hoffmann , Andreas Theodorou

AI Alignment is often presented as an interaction between a single designer and an artificial agent in which the designer attempts to ensure the agent's behavior is consistent with its purpose, and risks arise solely because of conflicts…

人工智能 · 计算机科学 2023-09-14 Steve Phelps , Rebecca Ranson

Large language model-based agents are rapidly evolving from simple conversational assistants into autonomous systems capable of performing complex, professional-level tasks in various domains. While these advancements promise significant…

This paper addresses the question of how to align AI systems with human values and situates it within a wider body of thought regarding technology and value. Far from existing in a vacuum, there has long been an interest in the ability of…

计算机与社会 · 计算机科学 2021-01-19 Iason Gabriel , Vafa Ghazavi

The deployment and use of AI systems should be both safe and broadly ethically acceptable. The principles-based ethics assurance argument pattern is one proposal in the AI ethics landscape that seeks to support and achieve that aim. The…

计算机与社会 · 计算机科学 2023-11-21 Marten H. L. Kaas , Zoe Porter , Ernest Lim , Aisling Higham , Sarah Khavandi , Ibrahim Habli

Human oversight of AI is promoted as a safeguard against risks such as inaccurate outputs, system malfunctions, or violations of fundamental rights, and is mandated in regulation like the European AI Act. Yet debates on human oversight have…

密码学与安全 · 计算机科学 2026-03-06 Jonas C. Ditz , Veronika Lazar , Elmar Lichtmeß , Carola Plesch , Matthias Heck , Kevin Baum , Markus Langer

The capabilities of artificial intelligence systems have been advancing to a great extent, but these systems still struggle with failure modes, vulnerabilities, and biases. In this paper, we study the current state of the field, and present…

密码学与安全 · 计算机科学 2025-06-12 Xingli Fang , Jianwei Li , Varun Mulchandani , Jung-Eun Kim

As machine learning systems become more powerful they also become increasingly unpredictable and opaque. Yet, finding human-understandable explanations of how they work is essential for their safe deployment. This technical report…

Training powerful AI systems to exhibit desired behaviors hinges on the ability to provide accurate human supervision on increasingly complex tasks. A promising approach to this problem is to amplify human judgement by leveraging the power…

人工智能 · 计算机科学 2025-06-17 Jonah Brown-Cohen , Geoffrey Irving , Georgios Piliouras

What makes safety claims about general purpose AI systems such as large language models trustworthy? We show that rather than the capabilities of security tools such as alignment and red teaming procedures, it is security practices based on…

密码学与安全 · 计算机科学 2025-07-30 Petr Spelda , Vit Stritecky

How can we ensure that AI systems are aligned with human values and remain safe? We can study this problem through the frameworks of the AI assistance and the AI shutdown games. The AI assistance problem concerns designing an AI agent that…

人工智能 · 计算机科学 2025-12-30 Alessio Benavoli , Alessandro Facchini , Marco Zaffalon

In the face of rapidly advancing AI technology, individuals will increasingly rely on AI agents to navigate life's growing complexities, raising critical concerns about maintaining both human agency and autonomy. This paper addresses a…

计算机与社会 · 计算机科学 2025-04-29 Philipp Koralus

This position paper states that AI Alignment in Multi-Agent Systems (MAS) should be considered a dynamic and interaction-dependent process that heavily depends on the social environment where agents are deployed, either collaborative,…

人工智能 · 计算机科学 2025-06-09 Florian Carichon , Aditi Khandelwal , Marylou Fauchard , Golnoosh Farnadi

AI alignment research aims to develop techniques to ensure that AI systems do not cause harm. However, every alignment technique has failure modes, which are conditions in which there is a non-negligible chance that the technique fails to…

人工智能 · 计算机科学 2025-10-14 Leonard Dung , Florian Mai

Artificial intelligence (AI) is now ubiquitous in our lives, and we regularly experience its decisions. Yet, the general public has very little knowledge about how it works, its use of data, its lack of objectivity, and its fallibility. In…

计算机与社会 · 计算机科学 2026-03-17 Carole Adam , Cedric Lauradoux

Recent developments in artificial intelligence (AI) have permeated through an array of different immersive environments, including virtual, augmented, and mixed realities. AI brings a wealth of potential that centers on its ability to…

人机交互 · 计算机科学 2024-05-10 Wangfan Li , Rohit Mallick , Carlos Toxtli-Hernandez , Christopher Flathmann , Nathan J. McNeese

In many contexts, lying -- the use of verbal falsehoods to deceive -- is harmful. While lying has traditionally been a human affair, AI systems that make sophisticated verbal statements are becoming increasingly prevalent. This raises the…

计算机与社会 · 计算机科学 2021-10-14 Owain Evans , Owen Cotton-Barratt , Lukas Finnveden , Adam Bales , Avital Balwit , Peter Wills , Luca Righetti , William Saunders

With the introduction of Artificial Intelligence (AI) and related technologies in our daily lives, fear and anxiety about their misuse as well as the hidden biases in their creation have led to a demand for regulation to address such…

人工智能 · 计算机科学 2021-04-09 The Anh Han , Tom Lenaerts , Francisco C. Santos , Luis Moniz Pereira

AI alignment aims to make AI systems behave in line with human intentions and values. As AI systems grow more capable, so do risks from misalignment. To provide a comprehensive and up-to-date overview of the alignment field, in this survey,…

AI alignment work is important from both a commercial and a safety lens. With this paper, we aim to help actors who support alignment efforts to make these efforts as effective as possible, and to avoid potential adverse effects. We begin…

计算机与社会 · 计算机科学 2023-12-18 Oliver Guest , Michael Aird , Seán Ó hÉigeartaigh