English
Related papers

Related papers: An alignment safety case sketch based on debate

200 papers

Artificial Intelligence (AI) is rapidly being integrated into critical systems across various domains, from healthcare to autonomous vehicles. While its integration brings immense benefits, it also introduces significant risks, including…

Computers and Society · Computer Science 2025-06-25 Zhiqiang Lin , Huan Sun , Ness Shroff

Rapid technological advancements in AI as well as the growing deployment of intelligent technologies in new application domains are currently driving the competition between businesses, nations and regions. This race for technological…

Computers and Society · Computer Science 2020-01-17 The Anh Han , Luis Moniz Pereira , Francisco C. Santos , Tom Lenaerts

Artificial intelligence (AI) technologies (re-)shape modern life, driving innovation in a wide range of sectors. However, some AI systems have yielded unexpected or undesirable outcomes or have been used in questionable manners. As a…

The field of Artificial Intelligence (AI) is going through a period of great expectations, introducing a certain level of anxiety in research, business and also policy. This anxiety is further energised by an AI race narrative that makes…

Artificial Intelligence · Computer Science 2021-06-09 The Anh Han , Luis Moniz Pereira , Tom Lenaerts , Francisco C. Santos

When AI agents don't align their actions with human values they may cause serious harm. One way to solve the value alignment problem is by including a human operator who monitors all of the agent's actions. Despite the fact, that this…

Human-Computer Interaction · Computer Science 2023-06-13 Yitzhak Spielberg , Amos Azaria

This paper presents a conceptual and operational framework for developing and operating safe and trustworthy AI agents based on a Three-Pillar Model grounded in transparency, accountability, and trustworthiness. Building on prior work in…

Computers and Society · Computer Science 2026-01-13 Edward C. Cheng , Jeshua Cheng , Alice Siu

In AI alignment, extensive latitude must be granted to annotators, either human or algorithmic, to judge which model outputs are `better' or `safer.' We refer to this latitude as alignment discretion. Such discretion remains largely…

Artificial Intelligence · Computer Science 2025-02-18 Maarten Buyl , Hadi Khalaf , Claudio Mayrink Verdun , Lucas Monteiro Paes , Caio C. Vieira Machado , Flavio du Pin Calmon

Alignment research focuses on making individual AI systems reliable. Human institutions achieve reliable collective behaviour differently: they mitigate the risk posed by misaligned individuals through organisational structure. Multi-agent…

Artificial Intelligence · Computer Science 2026-02-17 William Waites

Risk thresholds provide a measure of the level of risk exposure that a society or individual is willing to withstand, ultimately shaping how we determine the safety of technological systems. Against the backdrop of the Cold War, the first…

Computers and Society · Computer Science 2025-04-22 Heidy Khlaaf , Sarah Myers West

The leading AI companies are increasingly focused on building generalist AI agents -- systems that can autonomously plan, act, and pursue goals across almost all tasks that humans can perform. Despite how useful these systems might be,…

The issues of AI risk and AI safety are becoming critical as the prospect of artificial general intelligence (AGI) looms larger. The emergence of extremely large and capable generative models has led to alarming predictions and created a…

Artificial Intelligence · Computer Science 2025-05-20 Ali A. Minai

Artificial Intelligence (AI) is being increasingly used to develop systems that produce intelligent solutions. However, there is a major concern that whether the systems built will be trusted by humans. In order to establish trust in AI…

Artificial Intelligence · Computer Science 2020-05-13 Quratul-ain Mahesar , Simon Parsons

Effective collaboration between humans and AI-based systems requires effective modeling of the human in the loop, both in terms of the mental state as well as the physical capabilities of the latter. However, these models can also open up…

Artificial Intelligence · Computer Science 2018-01-31 Tathagata Chakraborti , Subbarao Kambhampati

The rapid advancement of artificial intelligence (AI) technologies presents profound challenges to societal safety. As AI systems become more capable, accessible, and integrated into critical services, the dual nature of their potential is…

Artificial Intelligence · Computer Science 2024-12-06 Giulio Corsi , Kyle Kilian , Richard Mallah

Frontier AI systems are increasingly capable and deployed in high-stakes multi-agent environments. However, existing AI safety benchmarks largely evaluate single agents, leaving multi-agent risks such as coordination failure and conflict…

Artificial Intelligence · Computer Science 2026-05-25 Pepijn Cobben , Xuanqiang Angelo Huang , Thao Amelia Pham , Isabel Dahlgren , Terry Jingchen Zhang , Zhijing Jin

Understanding the decisions made and actions taken by increasingly complex AI system remains a key challenge. This has led to an expanding field of research in explainable artificial intelligence (XAI), highlighting the potential of…

The neutrality thesis holds that technology cannot be laden with values. This long-standing view has faced critiques, but much of the argumentation against neutrality has focused on traditional, non-smart technologies like bridges and…

Artificial Intelligence · Computer Science 2024-08-23 Torben Swoboda , Lode Lauwaert

The progress of AI systems such as large language models (LLMs) raises increasingly pressing concerns about their safe deployment. This paper examines the value alignment problem for LLMs, arguing that current alignment strategies are…

Computation and Language · Computer Science 2025-06-06 Raphaël Millière

As AI agents surpass human capabilities, scalable oversight -- the problem of effectively supplying human feedback to potentially superhuman AI models -- becomes increasingly critical to ensure alignment. While numerous scalable oversight…

Artificial Intelligence · Computer Science 2025-04-08 Abhimanyu Pallavi Sudhir , Jackson Kaunismaa , Arjun Panickssery

In the future, artificial learning agents are likely to become increasingly widespread in our society. They will interact with both other learning agents and humans in a variety of complex settings including social dilemmas. We argue that…

Artificial Intelligence · Computer Science 2022-02-22 Tobias Baumann