中文
相关论文

相关论文: An alignment safety case sketch based on debate

200 篇论文

Frontier artificial intelligence (AI) systems pose increasing risks to society, making it essential for developers to provide assurances about their safety. One approach to offering such assurances is through a safety case: a structured,…

计算机与社会 · 计算机科学 2024-11-14 Arthur Goemans , Marie Davidsen Buhl , Jonas Schuett , Tomek Korbak , Jessica Wang , Benjamin Hilton , Geoffrey Irving

As LLM agents gain a greater capacity to cause harm, AI developers might increasingly rely on control measures such as monitoring to justify that they are safe. We sketch how developers could construct a "control safety case", which is a…

人工智能 · 计算机科学 2025-01-30 Tomek Korbak , Joshua Clymer , Benjamin Hilton , Buck Shlegeris , Geoffrey Irving

This position paper argues for two claims regarding AI testing and evaluation. First, to remain informative about deployment behaviour, evaluations need account for the possibility that AI systems understand their circumstances and reason…

计算机科学与博弈论 · 计算机科学 2025-08-22 Vojtech Kovarik , Eric Olav Chen , Sami Petersen , Alexis Ghersengorin , Vincent Conitzer

An assurance case is a structured argument, typically produced by safety engineers, to communicate confidence that a critical or complex system, such as an aircraft, will be acceptably safe within its intended context. Assurance cases often…

计算机与社会 · 计算机科学 2023-06-07 Zoe Porter , Ibrahim Habli , John McDermid , Marten Kaas

The ability to create artificial intelligence (AI) capable of performing complex tasks is rapidly outpacing our ability to ensure the safe and assured operation of AI-enabled systems. Fortunately, a landscape of AI safety research is…

Safety cases, structured arguments that a system is acceptably safe, are becoming central to the governance of AI systems. Yet, traditional safety-case practices from aviation or nuclear engineering rely on well-specified system boundaries,…

软件工程 · 计算机科学 2026-03-09 Sung Une Lee , Liming Zhu , Md Shamsujjoha , Liming Dong , Qinghua Lu , Jieshan Chen , Lionel Briand

As AI systems advance in capabilities, measuring their safety and alignment to human values is becoming paramount. A fast-growing field of AI research is devoted to developing such assessments. However, most current advances therein may be…

计算机与社会 · 计算机科学 2026-03-17 Max Hellrigel-Holderbaum , Edward James Young

As increasingly capable agents are deployed, a central safety challenge is how to retain meaningful human control without modifying the underlying system. We study a minimal control interface in which an agent chooses whether to act…

人工智能 · 计算机科学 2026-02-23 William Overman , Mohsen Bayati

Agentic AIs $-$ AIs that are capable and permitted to undertake complex actions with little supervision $-$ mark a new frontier in AI capabilities and raise new questions about how to safely create and align such systems with users,…

计算机与社会 · 计算机科学 2024-10-04 Hayley Clatterbuck , Clinton Castro , Arvo Muñoz Morán

While much research in artificial intelligence (AI) has focused on scaling capabilities, the accelerating pace of development makes countervailing work on producing harmless, "aligned" systems increasingly urgent. Yet research on alignment…

人工智能 · 计算机科学 2025-12-12 Dani Roytburg , Beck Miller

Creating systems that are aligned with our goals is seen as a leading approach to create safe and beneficial AI in both leading AI companies and the academic field of AI safety. We defend the view that misaligned AGI - future, generally…

计算机与社会 · 计算机科学 2025-06-05 Max Hellrigel-Holderbaum , Leonard Dung

A morally acceptable course of AI development should avoid two dangers: creating unaligned AI systems that pose a threat to humanity and mistreating AI systems that merit moral consideration in their own right. This paper argues these two…

计算机与社会 · 计算机科学 2025-10-16 Adam Bradley , Bradford Saad

Existing alignment research is dominated by concerns about safety and preventing harm: safeguards, controllability, and compliance. This paradigm of alignment parallels early psychology's focus on mental illness: necessary but incomplete.…

We introduce RedDebate, a novel multi-agent debate framework that provides the foundation for Large Language Models (LLMs) to identify and mitigate their unsafe behaviours. Existing AI safety approaches often rely on costly human evaluation…

计算与语言 · 计算机科学 2025-10-13 Ali Asad , Stephen Obadinma , Radin Shayanfar , Xiaodan Zhu

As AI systems grow more capable and autonomous, ensuring their safety and reliability requires not only model-level alignment but also strategic oversight of the humans and institutions involved in their development and deployment. Existing…

人工智能 · 计算机科学 2026-02-10 Cheol Woo Kim , Davin Choo , Tzeh Yuan Neoh , Milind Tambe

As frontier artificial intelligence (AI) systems become more capable, it becomes more important that developers can explain why their systems are sufficiently safe. One way to do so is via safety cases: reports that make a structured…

计算机与社会 · 计算机科学 2024-10-30 Marie Davidsen Buhl , Gaurav Sett , Leonie Koessler , Jonas Schuett , Markus Anderljung

The range of application of artificial intelligence (AI) is vast, as is the potential for harm. Growing awareness of potential risks from AI systems has spurred action to address those risks, while eroding confidence in AI systems and the…

Scalable oversight protocols aim to enable humans to accurately supervise superhuman AI. In this paper we study debate, where two AI's compete to convince a judge; consultancy, where a single AI tries to convince a judge that asks…

Evaluating the safety of AI Systems is a pressing concern for organizations deploying them. In addition to the societal damage done by the lack of fairness of those systems, deployers are concerned about the legal repercussions and the…

A leading proposal for aligning artificial superintelligence (ASI) is to use AI agents to automate an increasing fraction of alignment research as capabilities improve. We argue that, even when research agents are not scheming to…

人工智能 · 计算机科学 2026-05-18 Aleksandr Bowkis , Marie Davidsen Buhl , Jacob Pfau , Geoffrey Irving