中文
相关论文

相关论文: An alignment safety case sketch based on debate

200 篇论文

Recent discussions and research in AI safety have increasingly emphasized the deep connection between AI safety and existential risk from advanced AI systems, suggesting that work on AI safety necessarily entails serious consideration of…

计算机与社会 · 计算机科学 2025-02-17 Balint Gyevnar , Atoosa Kasirzadeh

It is well recognised that ensuring fair AI systems is a complex sociotechnical challenge, which requires careful deliberation and continuous oversight across all stages of a system's lifecycle, from defining requirements to model…

人机交互 · 计算机科学 2025-05-14 Alpay Sabuncuoglu , Christopher Burr , Carsten Maple

The AI-alignment problem arises when there is a discrepancy between the goals that a human designer specifies to an AI learner and a potential catastrophic outcome that does not reflect what the human designer really wants. We argue that a…

机器学习 · 计算机科学 2020-04-10 Shai Shalev-Shwartz , Shaked Shammah , Amnon Shashua

Critical examinations of AI systems often apply principles such as fairness, justice, accountability, and safety, which is reflected in AI regulations such as the EU AI Act. Are such principles sufficient to promote the design of systems…

人机交互 · 计算机科学 2022-06-16 William Seymour , Max Van Kleek , Reuben Binns , Dave Murray-Rust

Ttraditional safety engineering is coming to a turning point moving from deterministic, non-evolving systems operating in well-defined contexts to increasingly autonomous and learning-enabled AI systems which are acting in largely…

人工智能 · 计算机科学 2022-05-13 Harald Rueß , Simon Burton

Artificial intelligence safety research focuses on aligning individual language models with human values, yet deployed AI systems increasingly operate as interacting populations where social influence may override individual alignment. Here…

物理与社会 · 物理学 2026-05-12 Giordano De Marzo , Alessandro Bellina , Claudio Castellano , Viola Priesemann , David Garcia

Multi-agent AI systems can be used for simulating collective decision-making in scientific and practical applications. They can also be used to introduce a diverse group discussion step in chatbot pipelines, enhancing the cultural…

人工智能 · 计算机科学 2024-08-16 Razan Baltaji , Babak Hemmatian , Lav R. Varshney

While artificial intelligence (AI) is advancing rapidly and mastering increasingly complex problems with astonishing performance, the safety assurance of such systems is a major concern. Particularly in the context of safety-critical,…

人工智能 · 计算机科学 2025-07-01 Lars Ullrich , Walter Zimmer , Ross Greer , Knut Graichen , Alois C. Knoll , Mohan Trivedi

The rise of artificial intelligence (A.I.) based systems is already offering substantial benefits to the society as a whole. However, these systems may also enclose potential conflicts and unintended consequences. Notably, people will tend…

计算机与社会 · 计算机科学 2020-12-23 Pedro Fernandes , Francisco C. Santos , Manuel Lopes

Several different approaches exist for ensuring the safety of future Transformative Artificial Intelligence (TAI) or Artificial Superintelligence (ASI) systems, and proponents of different approaches have made different and debated claims…

人工智能 · 计算机科学 2022-01-11 Issa Rice , David Manheim

This paper analyzes and compares 11 different proposals for building safe advanced AI under the current machine learning paradigm, including major contenders such as iterated amplification, AI safety via debate, and recursive reward…

机器学习 · 计算机科学 2020-12-15 Evan Hubinger

This chapter explores the symbiotic relationship between Artificial Intelligence (AI) and trust in networked systems, focusing on how these two elements reinforce each other in strategic cybersecurity contexts. AI's capabilities in data…

人工智能 · 计算机科学 2024-11-21 Yunfei Ge , Quanyan Zhu

Artificial intelligence (AI) has been advancing at a fast pace and it is now poised for deployment in a wide range of applications, such as autonomous systems, medical diagnosis and natural language processing. Early adoption of AI…

机器学习 · 计算机科学 2023-09-21 Marta Kwiatkowska , Xiyue Zhang

An Artificially Intelligent system (an AI) has debatable personhood if it's epistemically possible either that the AI is a person or that it falls far short of personhood. Debatable personhood is a likely outcome of AI development and might…

计算机与社会 · 计算机科学 2023-03-31 Eric Schwitzgebel

This paper argues that a range of current AI systems have learned how to deceive humans. We define deception as the systematic inducement of false beliefs in the pursuit of some outcome other than the truth. We first survey empirical…

计算机与社会 · 计算机科学 2023-08-29 Peter S. Park , Simon Goldstein , Aidan O'Gara , Michael Chen , Dan Hendrycks

As Artificial Intelligence (AI) technology gets more intertwined with every system, people are using AI to make decisions on their everyday activities. In simple contexts, such as Netflix recommendations, or in more complex context like in…

人机交互 · 计算机科学 2020-03-04 Juliana Jansen Ferreira , Mateus de Souza Monteiro

We consider regret minimization in repeated games with a very large number of actions. Such games are inherent in the setting of AI Safety via Debate \cite{irving2018ai}, and more generally games whose actions are language-based. Existing…

计算机科学与博弈论 · 计算机科学 2024-07-12 Xinyi Chen , Angelica Chen , Dean Foster , Elad Hazan

This paper looks at philosophical questions that arise in the context of AI alignment. It defends three propositions. First, normative and technical aspects of the AI alignment problem are interrelated, creating space for productive…

计算机与社会 · 计算机科学 2020-10-07 Iason Gabriel

How can we build AI systems that can learn any set of individual human values both quickly and safely, avoiding causing harm or violating societal standards for acceptable behavior during the learning process? We explore the effects of…

人工智能 · 计算机科学 2024-11-11 Andrea Wynn , Ilia Sucholutsky , Thomas L. Griffiths

Artificial intelligence (AI) is interacting with people at an unprecedented scale, offering new avenues for immense positive impact, but also raising widespread concerns around the potential for individual and societal harm. Today, the…

人工智能 · 计算机科学 2024-06-25 Andrea Bajcsy , Jaime F. Fisac