English
Related papers

Related papers: A Decision-Theoretic Approach for Managing Misalig…

200 papers

Deployed, autonomous AI systems must often evaluate multiple plausible courses of action (extended sequences of behavior) in novel or under-specified contexts. Despite extensive training, these systems will inevitably encounter scenarios…

Artificial Intelligence · Computer Science 2025-11-19 Steven J. Jones , Robert E. Wray , John E. Laird

We suggest that the analysis of incomplete contracting developed by law and economics researchers can provide a useful framework for understanding the AI alignment problem and help to generate a systematic approach to finding solutions. We…

Artificial Intelligence · Computer Science 2018-04-13 Dylan Hadfield-Menell , Gillian Hadfield

When machine learning is outsourced to a rational agent, conflicts of interest might arise and severely impact predictive performance. In this work, we propose a theoretical framework for incentive-aware delegation of machine learning…

Machine Learning · Computer Science 2023-12-07 Eden Saig , Inbal Talgam-Cohen , Nir Rosenfeld

As artificial intelligence (AI) becomes increasingly integrated into workflows, humans must decide when to rely on AI advice. These decisions depend on general efficacy beliefs, i.e., humans' confidence in their own abilities and their…

Human-Computer Interaction · Computer Science 2026-03-12 Philipp Spitzer , Joshua Holstein

The AI alignment problem, which focusses on ensuring that artificial intelligence (AI), including AGI and ASI, systems act according to human values, presents profound challenges. With the progression from narrow AI to Artificial General…

Artificial Intelligence · Computer Science 2025-07-25 Alberto Hernández-Espinosa , Felipe S. Abrahão , Olaf Witkowski , Hector Zenil

In this paper we propose an advanced approach to integrating artificial intelligence (AI) into healthcare: autonomous decision support. This approach allows the AI algorithm to act autonomously for a subset of patient cases whilst serving a…

Artificial Intelligence · Computer Science 2025-03-25 Yan Jia , Harriet Evans , Zoe Porter , Simon Graham , John McDermid , Tom Lawton , David Snead , Ibrahim Habli

Collaboration with artificial intelligence (AI) has improved human decision-making across various domains by leveraging the complementary capabilities of humans and AI. Yet, humans systematically overrely on AI advice, even when their…

Human-Computer Interaction · Computer Science 2026-05-15 Joshua Holstein , Patrick Hemmer , Gerhard Satzger , Wei Sun

This position paper argues for metacognition as a general design principle for creating more accurate, secure, and efficient AI. The metacognitive solution involves systems monitoring their own states and judiciously allocating resources…

Artificial Intelligence · Computer Science 2026-05-18 Sergei Chuprov , Richard D. Lange , Leon Reznik , Paulo Shakarian , Raman Zatsarenko , Dmitrii Korobeinikov

AI alignment is about ensuring AI systems only pursue goals and activities that are beneficial to humans. Most of the current approach to AI alignment is to learn what humans value from their behavioural data. This paper proposes a…

Artificial Intelligence · Computer Science 2023-10-06 Pei-Yu Chen , Myrthe L. Tielman , Dirk K. J. Heylen , Catholijn M. Jonker , M. Birna van Riemsdijk

AI-related incidents are becoming increasingly frequent and severe, ranging from safety failures to misuse by malicious actors. In such complex situations, identifying which elements caused an adverse outcome, the problem of cause…

Artificial Intelligence · Computer Science 2026-03-17 Maria Victoria Carro , David Lagnado

There are many settings in which a principal performs a task by delegating it to an agent, who searches over possible solutions and proposes one to the principal. This describes many aspects of the workflow within organizations, as well as…

Computer Science and Game Theory · Computer Science 2018-06-20 Jon Kleinberg , Robert Kleinberg

This paper looks at philosophical questions that arise in the context of AI alignment. It defends three propositions. First, normative and technical aspects of the AI alignment problem are interrelated, creating space for productive…

Computers and Society · Computer Science 2020-10-07 Iason Gabriel

In high-stakes AI-supported decisions, considerations are not purely technical but involve moral judgments about fairness, responsibility, and harm. While prior research has focused mainly on functional or behavioral alignment, this paper…

Human-Computer Interaction · Computer Science 2026-04-17 Christiane Ernst , Luis Gutmann , Domenique Zipperling , Kathrin Figl , Niklas Kühl

Appropriate reliance is critical to achieving synergistic human-AI collaboration. For instance, when users over-rely on AI assistance, their human-AI team performance is bounded by the model's capability. This work studies how the…

Human-Computer Interaction · Computer Science 2024-01-17 Shiye Cao , Anqi Liu , Chien-Ming Huang

When robots share the same workspace with other intelligent agents (e.g., other robots or humans), they must be able to reason about the behaviors of their neighboring agents while accomplishing the designated tasks. In practice,…

Robotics · Computer Science 2022-10-18 Junhong Xu , Durgakant Pushp , Kai Yin , Lantao Liu

Creating systems that are aligned with our goals is seen as a leading approach to create safe and beneficial AI in both leading AI companies and the academic field of AI safety. We defend the view that misaligned AGI - future, generally…

Computers and Society · Computer Science 2025-06-05 Max Hellrigel-Holderbaum , Leonard Dung

Many settings of interest involving humans and machines -- from virtual personal assistants to autonomous vehicles -- can naturally be modelled as principals (humans) delegating to agents (machines), which then interact with each other on…

Computer Science and Game Theory · Computer Science 2024-08-07 Oliver Sourbut , Lewis Hammond , Harriet Wood

AI-driven decision-making systems are becoming instrumental in the public sector, with applications spanning areas like criminal justice, social welfare, financial fraud detection, and public health. While these systems offer great…

Machine Learning · Computer Science 2024-10-15 Unai Fischer-Abaigar , Christoph Kern , Noam Barda , Frauke Kreuter

Jurisprudence, the study of how judges should properly decide cases, and alignment, the science of getting AI models to conform to human values, share a fundamental structure. These seemingly distant fields both seek to predict and shape…

Artificial Intelligence · Computer Science 2026-05-12 Nicholas Caputo

Value alignment has emerged in recent years as a basic principle to produce beneficial and mindful Artificial Intelligence systems. It mainly states that autonomous entities should behave in a way that is aligned with our human values. In…

Multiagent Systems · Computer Science 2021-06-28 Nieves Montes , Carles Sierra