English
Related papers

Related papers: Misalignment or misuse? The AGI alignment tradeoff

200 papers

The rapid advancement of artificial intelligence (AI) systems suggests that artificial general intelligence (AGI) systems may soon arrive. Many researchers are concerned that AIs and AGIs will harm humans via intentional misuse (AI-misuse)…

Artificial Intelligence · Computer Science 2023-05-31 Catalin Mitelut , Ben Smith , Peter Vamplew

Existing alignment research is dominated by concerns about safety and preventing harm: safeguards, controllability, and compliance. This paradigm of alignment parallels early psychology's focus on mental illness: necessary but incomplete.…

As AI adoption expands across human society, the problem of aligning AI models to match human preferences remains a grand challenge. Currently, the AI alignment field is deeply divided between behavioral and representational approaches,…

Computers and Society · Computer Science 2025-08-12 Ben Y. Reis , William La Cava

Artificial Intelligence (AI) is one of the most transformative technologies of the 21st century. The extent and scope of future AI capabilities remain a key uncertainty, with widespread disagreement on timelines and potential impacts. As…

Artificial Intelligence · Computer Science 2023-11-27 Kyle A. Kilian , Christopher J. Ventura , Mark M. Bailey

As artificial intelligence scales, the concepts of alignment, agency, and autonomy have become central to AI safety, governance, and control. However, even in human contexts, these terms lack universal definitions, varying across…

Computers and Society · Computer Science 2025-03-11 Krti Tallam

As artificial intelligence (AI) becomes deeply integrated into critical infrastructures and everyday life, ensuring its safe deployment is one of humanity's most urgent challenges. Current AI models prioritize task optimization over safety,…

Artificial Intelligence · Computer Science 2024-11-08 Joshua T. S. Hewson

As AI systems become embedded in everyday practice, value misalignment has emerged as a pressing concern. Yet, dominant alignment approaches remain model centric, treating users as passive recipients of prespecified values rather than as…

Human-Computer Interaction · Computer Science 2026-04-22 Anne Arzberger , Enrico Liscio , Maria Luce Lupetti , Inigo Martinez de Rituerto de Troya , Jie Yang

Artificial intelligence (AI) is advancing exponentially and is likely to have profound impacts on human wellbeing, social equity, and environmental sustainability. Here we argue that the "alignment problem" in AI research is also an…

General Economics · Economics 2026-04-30 Daniel W. O'Neill , Stefano Vrizzi , Noemi Luna Carmeno , Felix Creutzig , Jefim Vogel

While much research in artificial intelligence (AI) has focused on scaling capabilities, the accelerating pace of development makes countervailing work on producing harmless, "aligned" systems increasingly urgent. Yet research on alignment…

Artificial Intelligence · Computer Science 2025-12-12 Dani Roytburg , Beck Miller

AI-based systems have been used widely across various industries for different decisions ranging from operational decisions to tactical and strategic ones in low- and high-stakes contexts. Gradually the weaknesses and issues of these…

Human-Computer Interaction · Computer Science 2022-01-13 Morteza Saberi

Artificial intelligence (AI) systems will increasingly be used to cause harm as they grow more capable. In fact, AI systems are already starting to be used to automate fraudulent activities, violate human rights, create harmful fake images,…

Artificial Intelligence · Computer Science 2023-03-30 Markus Anderljung , Julian Hazell

Existing work on the alignment problem has focused mainly on (1) qualitative descriptions of the alignment problem; (2) attempting to align AI actions with human interests by focusing on value specification and learning; and/or (3) focusing…

Multiagent Systems · Computer Science 2025-06-03 Aidan Kierans , Avijit Ghosh , Hananel Hazan , Shiri Dori-Hacohen

This memorandum presents four recommendations aimed at strengthening the principles of AI model reliability and AI model governability, as DoW, ODNI, NIST, and CAISI refine AI assurance frameworks under the AI Action Plan. Our focus…

Computers and Society · Computer Science 2025-10-13 Matteo Pistillo , Charlotte Stix

Recent progress in artificial intelligence (AI) has drawn attention to the technology's transformative potential, including what some see as its prospects for causing large-scale harm. We review two influential arguments purporting to show…

Computers and Society · Computer Science 2024-01-30 Adam Bales , William D'Alessandro , Cameron Domenico Kirk-Giannini

While autonomous agents often surpass humans in their ability to handle vast and complex data, their potential misalignment (i.e., lack of transparency regarding their true objective) has thus far hindered their use in critical applications…

Artificial Intelligence · Computer Science 2024-12-03 Frédéric Berdoz , Roger Wattenhofer

As Artificial Intelligence (AI) advances toward Artificial General Intelligence (AGI) and eventually Artificial Superintelligence (ASI), it may potentially surpass human control, deviate from human values, and even lead to irreversible…

The expanding application of Artificial Intelligence (AI) in scientific fields presents unprecedented opportunities for discovery and innovation. However, this growth is not without risks. AI models in science, if misused, can amplify risks…

Artificial Intelligence · Computer Science 2023-12-12 Jiyan He , Weitao Feng , Yaosen Min , Jingwei Yi , Kunsheng Tang , Shuai Li , Jie Zhang , Kejiang Chen , Wenbo Zhou , Xing Xie , Weiming Zhang , Nenghai Yu , Shuxin Zheng

There is a substantial and ever-growing corpus of evidence and literature exploring the impacts of Artificial intelligence (AI) technologies on society, politics, and humanity as a whole. A separate, parallel body of work has explored…

Computers and Society · Computer Science 2022-09-23 Benjamin S. Bucknall , Shiri Dori-Hacohen

Existing evaluations of AI misuse safeguards provide a patchwork of evidence that is often difficult to connect to real-world decisions. To bridge this gap, we describe an end-to-end argument (a "safety case") that misuse safeguards reduce…

Machine Learning · Computer Science 2025-05-26 Joshua Clymer , Jonah Weinbaum , Robert Kirk , Kimberly Mai , Selena Zhang , Xander Davies

Artificial Intelligence (AI) is rapidly being integrated into critical systems across various domains, from healthcare to autonomous vehicles. While its integration brings immense benefits, it also introduces significant risks, including…

Computers and Society · Computer Science 2025-06-25 Zhiqiang Lin , Huan Sun , Ness Shroff