English
Related papers

Related papers: Research Superalignment Should Advance Now with Al…

200 papers

As machine learning (ML) systems have advanced, they have acquired more power over humans' lives, and questions about what values are embedded in them have become more complex and fraught. It is conceivable that in the coming decades,…

Computers and Society · Computer Science 2019-03-18 Sky Croeser , Peter Eckersley

While autonomous agents often surpass humans in their ability to handle vast and complex data, their potential misalignment (i.e., lack of transparency regarding their true objective) has thus far hindered their use in critical applications…

Artificial Intelligence · Computer Science 2024-12-03 Frédéric Berdoz , Roger Wattenhofer

Artificial intelligence (AI) has made significant strides in recent years, yet it continues to struggle with a fundamental aspect of cognition present in all animals: common sense. Current AI systems, including those designed for complex…

Artificial Intelligence · Computer Science 2025-01-14 Hugo Latapie

As progress in AI continues to advance, it is important to know how advanced systems will make choices and in what ways they may fail. Machines can already outsmart humans in some domains, and understanding how to safely build ones which…

Artificial Intelligence · Computer Science 2023-04-04 Stephen Casper

Artificial general intelligence (AGI) does not yet exist, but given the pace of technological development in artificial intelligence, it is projected to reach human-level intelligence within roughly the next two decades. After that, many…

Computers and Society · Computer Science 2023-11-16 David R. Mandel

We suggest that the analysis of incomplete contracting developed by law and economics researchers can provide a useful framework for understanding the AI alignment problem and help to generate a systematic approach to finding solutions. We…

Artificial Intelligence · Computer Science 2018-04-13 Dylan Hadfield-Menell , Gillian Hadfield

Aligning powerful AI models on tasks that surpass human evaluation capabilities is the central problem of \textbf{superalignment}. To address this problem, weak-to-strong generalization aims to elicit the capabilities of strong models…

Machine Learning · Computer Science 2025-03-07 Junhao Shi , Qinyuan Cheng , Zhaoye Fei , Yining Zheng , Qipeng Guo , Xipeng Qiu

For humanity to maintain and expand its agency into the future, the most powerful systems we create must be those which act to align the future with the will of humanity. The most powerful systems today are massive institutions like…

Value alignment problems arise in scenarios where the specified objectives of an AI agent don't match the true underlying objective of its users. The problem has been widely argued to be one of the central safety problems in AI.…

Artificial Intelligence · Computer Science 2023-02-10 Malek Mechergui , Sarath Sreedharan

The increasing prevalence of artificial agents creates a correspondingly increasing need to manage disagreements between humans and artificial agents, as well as between artificial agents themselves. Considering this larger space of…

Neurons and Cognition · Quantitative Biology 2023-10-23 Kerem Oktar , Ilia Sucholutsky , Tania Lombrozo , Thomas L. Griffiths

AI Safety researchers attempting to align values of highly capable intelligent systems with those of humanity face a number of challenges including personal value extraction, multi-agent value merger and finally in-silico encoding.…

Artificial Intelligence · Computer Science 2019-01-08 Roman V. Yampolskiy

Collaboration with artificial intelligence (AI) has improved human decision-making across various domains by leveraging the complementary capabilities of humans and AI. Yet, humans systematically overrely on AI advice, even when their…

Human-Computer Interaction · Computer Science 2026-05-15 Joshua Holstein , Patrick Hemmer , Gerhard Satzger , Wei Sun

Neuroscience and Artificial Intelligence (AI) have made impressive progress in recent years but remain only loosely interconnected. Based on a workshop convened by the National Science Foundation in August 2025, we identify three…

The complexity of psychological principles underscore a significant societal challenge, given the vast social implications of psychological problems. Bridging the gap between understanding these principles and their actual clinical and…

Artificial Intelligence · Computer Science 2023-12-11 Tianyu He , Guanghui Fu , Yijing Yu , Fan Wang , Jianqiang Li , Qing Zhao , Changwei Song , Hongzhi Qi , Dan Luo , Huijing Zou , Bing Xiang Yang

Artificial Intelligence (AI) is one of the most transformative technologies of the 21st century. The extent and scope of future AI capabilities remain a key uncertainty, with widespread disagreement on timelines and potential impacts. As…

Artificial Intelligence · Computer Science 2023-11-27 Kyle A. Kilian , Christopher J. Ventura , Mark M. Bailey

Artificial Intelligence (AI) is progressing rapidly, and companies are shifting their focus to developing generalist AI systems that can autonomously act and pursue goals. Increases in capabilities and autonomy may soon massively amplify…

We introduce an increasing-complexity, open-ended, and human-agnostic metric to evaluate foundational and frontier AI models in the context of Artificial General Intelligence (AGI) and Artificial Super Intelligence (ASI) claims. Unlike…

Artificial Intelligence · Computer Science 2026-02-13 Alberto Hernández-Espinosa , Luan Ozelim , Felipe S. Abrahão , Hector Zenil

General intelligence, the ability to solve arbitrary solvable problems, is supposed by many to be artificially constructible. Narrow intelligence, the ability to solve a given particularly difficult problem, has seen impressive recent…

Artificial Intelligence · Computer Science 2020-07-22 Michael K Cohen , Badri Vellambi , Marcus Hutter

Human feedback is critical for aligning AI systems to human values. As AI capabilities improve and AI is used to tackle more challenging tasks, verifying quality and safety becomes increasingly challenging. This paper explores how we can…

Artificial Intelligence · Computer Science 2025-10-31 Rishub Jain , Sophie Bridgers , Lili Janzer , Rory Greig , Tian Huey Teh , Vladimir Mikulik

The value-alignment problem for artificial intelligence (AI) asks how we can ensure that the 'values' (i.e., objective functions) of artificial systems are aligned with the values of humanity. In this paper, I argue that linguistic…

Artificial Intelligence · Computer Science 2022-07-05 Travis LaCroix
‹ Prev 1 3 4 5 6 7 10 Next ›