English
Related papers

Related papers: Alignment has a Fantasia Problem

200 papers

Human feedback is commonly utilized to finetune AI assistants. But human feedback may also encourage model responses that match user beliefs over truthful ones, a behaviour known as sycophancy. We investigate the prevalence of sycophancy in…

Machine Learning algorithms are technological key enablers for artificial intelligence (AI). Due to the inherent complexity, these learning algorithms represent black boxes and are difficult to comprehend, therefore influencing compliance…

Computers and Society · Computer Science 2020-02-21 NIklas Kuhl , Jodie Lobana , Christian Meske

As AI becomes more prevalent throughout society, effective methods of integrating humans and AI systems that leverage their respective strengths and mitigate risk have become an important priority. In this paper, we introduce the paradigm…

Machine Learning · Computer Science 2023-10-24 Jiayi Wang , Zhengling Qi , Chengchun Shi

The dominant practice of AI alignment assumes (1) that preferences are an adequate representation of human values, (2) that human rationality can be understood in terms of maximizing the satisfaction of preferences, and (3) that AI systems…

Artificial Intelligence · Computer Science 2024-11-12 Tan Zhi-Xuan , Micah Carroll , Matija Franklin , Hal Ashton

The project of aligning machine behavior with human values raises a basic problem: whose moral expectations should guide AI decision-making? Much alignment research assumes that the appropriate benchmark is how humans themselves would act…

Computers and Society · Computer Science 2026-05-13 Benjamin Minhao Chen , Xinyu Xie

According to several empirical investigations, despite enhancing human capabilities, human-AI cooperation frequently falls short of expectations and fails to reach true synergy. We propose a task-driven framework that reverses prevalent…

Computers and Society · Computer Science 2026-05-26 Saleh Afroogh , Kush R. Varshney , Jason D'Cruz

When working on digital devices, people often face distractions that can lead to a decline in productivity and efficiency, as well as negative psychological and emotional impacts. To address this challenge, we introduce a novel Artificial…

Human-Computer Interaction · Computer Science 2026-03-03 Juheon Choi , Juyong Lee , Jian Kim , Chanyoung Kim , Taywon Min , W. Bradley Knox , Min Kyung Lee , Kimin Lee

AI systems are being deployed to support human decision making in high-stakes domains. In many cases, the human and AI form a team, in which the human makes decisions after reviewing the AI's inferences. A successful partnership requires…

Human-Computer Interaction · Computer Science 2019-06-06 Gagan Bansal , Besmira Nushi , Ece Kamar , Dan Weld , Walter Lasecki , Eric Horvitz

As reliance on AI systems for decision-making grows, it becomes critical to ensure that human users can appropriately balance trust in AI suggestions with their own judgment, especially in high-stakes domains like healthcare. However, human…

Human-Computer Interaction · Computer Science 2025-01-29 Zichen Chen , Yunhao Luo , Misha Sra

Isolated perspectives have often paved the way for great scientific discoveries. However, many breakthroughs only emerged when moving away from singular views towards interactions. Discussions on Artificial Intelligence (AI) typically treat…

Human-Computer Interaction · Computer Science 2025-04-29 Nick von Felten

Alignment of artificial intelligence (AI) encompasses the normative problem of specifying how AI systems should act and the technical problem of ensuring AI systems comply with those specifications. To date, AI alignment has generally…

The increasing prevalence of artificial agents creates a correspondingly increasing need to manage disagreements between humans and artificial agents, as well as between artificial agents themselves. Considering this larger space of…

Neurons and Cognition · Quantitative Biology 2023-10-23 Kerem Oktar , Ilia Sucholutsky , Tania Lombrozo , Thomas L. Griffiths

As AI systems grow increasingly capable of operating for hours or days at a time, users' prompts are transforming into elaborate specifications for the AI to autonomously work on. While prompting for bounded, single-turn tasks has been…

Human-Computer Interaction · Computer Science 2026-04-07 Savvas Petridis , Michael Xieyang Liu , Alexander J. Fiannaca , Carrie J. Cai , Michael Terry

To assist users in complex tasks, LLMs generate plans: step-by-step instructions towards a goal. While alignment methods aim to ensure LLM plans are helpful, they train (RLHF) or evaluate (ChatbotArena) on what users prefer, assuming this…

As Artificial Intelligence (AI) technology becomes more and more prevalent, it becomes increasingly important to explore how we as humans interact with AI. The Human-AI Interaction (HAI) sub-field has emerged from the Human-Computer…

Human-Computer Interaction · Computer Science 2024-01-15 Mark Adkins

General Alignment has improved average-case helpfulness and safety, but current alignment practice still rewards confident, single-turn responses. The problem is not only that models fail on edge cases; it is that current evaluation makes…

Computation and Language · Computer Science 2026-05-19 Han Bao , Yue Huang , Xiaoda Wang , Zheyuan Zhang , Yujun Zhou , Carl Yang , Xiangliang Zhang , Yanfang Ye

Agentic AIs $-$ AIs that are capable and permitted to undertake complex actions with little supervision $-$ mark a new frontier in AI capabilities and raise new questions about how to safely create and align such systems with users,…

Computers and Society · Computer Science 2024-10-04 Hayley Clatterbuck , Clinton Castro , Arvo Muñoz Morán

When deployed, AI agents will encounter problems that are beyond their autonomous problem-solving capabilities. Leveraging human assistance can help agents overcome their inherent limitations and robustly cope with unfamiliar situations. We…

Machine Learning · Computer Science 2022-06-24 Khanh Nguyen , Yonatan Bisk , Hal Daumé

Appropriate Trust in Artificial Intelligence (AI) systems has rapidly become an important area of focus for both researchers and practitioners. Various approaches have been used to achieve it, such as confidence scores, explanations,…

Human-Computer Interaction · Computer Science 2023-11-14 Siddharth Mehrotra , Chadha Degachi , Oleksandra Vereschak , Catholijn M. Jonker , Myrthe L. Tielman

We describe a class of tasks called decision-oriented dialogues, in which AI assistants such as large language models (LMs) must collaborate with one or more humans via natural language to help them make complex decisions. We formalize…

Computation and Language · Computer Science 2024-05-07 Jessy Lin , Nicholas Tomlin , Jacob Andreas , Jason Eisner
‹ Prev 1 4 5 6 7 8 10 Next ›