中文
相关论文

相关论文: Towards Understanding Sycophancy in Language Model…

200 篇论文

Large language models are increasingly used in decision-making tasks that require them to process information from a variety of sources, including both human experts and other algorithmic agents. How do LLMs weigh the information provided…

人工智能 · 计算机科学 2026-02-26 Jessica Y. Bo , Lillio Mok , Ashton Anderson

Humans rely more and more on systems with AI components. The AI community typically treats human inputs as a given and optimizes AI models only. This thinking is one-sided and it neglects the fact that humans can learn, too. In this work,…

人机交互 · 计算机科学 2020-09-22 Johannes Schneider

Artificial Intelligence (AI) has been used extensively in automatic decision making in a broad variety of scenarios, ranging from credit ratings for loans to recommendations of movies. Traditional design guidelines for AI models focus…

人工智能 · 计算机科学 2018-09-27 Marisa Vasconcelos , Carlos Cardonha , Bernardo Gonçalves

Spurious correlations were found to be an important factor explaining model performance in various NLP tasks (e.g., gender or racial artifacts), often considered to be ''shortcuts'' to the actual task. However, humans tend to similarly make…

计算与语言 · 计算机科学 2025-08-25 Gili Lior , Gabriel Stanovsky

AI systems increasingly support human decision-making. In many cases, despite the algorithm's superior performance, the final decision remains in human hands. For example, an AI may assist doctors in determining which diagnostic tests to…

人工智能 · 计算机科学 2026-02-20 Gali Noti , Kate Donahue , Jon Kleinberg , Sigal Oren

Learning from human feedback is a prominent technique to align the output of large language models (LLMs) with human expectations. Reinforcement learning from human feedback (RLHF) leverages human preference signals that are in the form of…

计算与语言 · 计算机科学 2023-11-27 Di Jin , Shikib Mehri , Devamanyu Hazarika , Aishwarya Padmakumar , Sungjin Lee , Yang Liu , Mahdi Namazifar

Large language models (LLMs) have achieved strong performance across a wide range of tasks, but they are also prone to sycophancy, the tendency to agree with user statements regardless of validity. Previous research has outlined both the…

计算与语言 · 计算机科学 2026-03-31 Bayan Abdullah Aldahlawi , A. B. M. Ashikur Rahman , Irfan Ahmad

Conversational human-likeness plays a central role in human-AI interaction, yet it has remained difficult to define, measure, and optimize. As a result, improvements in human-like behavior are largely driven by scale or broad supervised…

人工智能 · 计算机科学 2026-01-08 Masum Hasan , Junjie Zhao , Ehsan Hoque

Learning from human feedback has become a pivot technique in aligning large language models (LLMs) with human preferences. However, acquiring vast and premium human feedback is bottlenecked by time, labor, and human capability, resulting in…

计算与语言 · 计算机科学 2024-07-17 Ganqu Cui , Lifan Yuan , Ning Ding , Guanming Yao , Bingxiang He , Wei Zhu , Yuan Ni , Guotong Xie , Ruobing Xie , Yankai Lin , Zhiyuan Liu , Maosong Sun

In many real world contexts, successful human-AI collaboration requires humans to productively integrate complementary sources of information into AI-informed decisions. However, in practice human decision-makers often lack understanding of…

人机交互 · 计算机科学 2023-01-30 Kenneth Holstein , Maria De-Arteaga , Lakshmi Tumati , Yanghuidi Cheng

We develop a decision-theoretic model of human-AI interaction to study when AI assistance improves or impairs human decision-making. A human decision-maker observes private information and receives a recommendation from an AI system, but…

计算机科学与博弈论 · 计算机科学 2026-02-17 Saurabh Amin , Amine Bennouna , Daniel Huttenlocher , Dingwen Kong , Liang Lyu , Asuman Ozdaglar

Some traits making a "good" AI model are hard to describe upfront. For example, should responses be more polite or more casual? Such traits are sometimes summarized as model character or personality. Without a clear objective, conventional…

计算与语言 · 计算机科学 2025-10-01 Arduin Findeis , Timo Kaufmann , Eyke Hüllermeier , Robert Mullins

As algorithmic tools increasingly aid experts in making consequential decisions, the need to understand the precise factors that mediate their influence has grown commensurately. In this paper, we present a crowdsourcing vignette study…

人机交互 · 计算机科学 2022-05-20 Riccardo Fogliato , Sina Fazelpour , Shantanu Gupta , Zachary Lipton , David Danks

Behavioral scientists have classically documented aversion to algorithmic decision aids, from simple linear models to AI. Sentiment, however, is changing and possibly accelerating AI helper usage. AI assistance is, arguably, most valuable…

人工智能 · 计算机科学 2023-07-28 Nikolos Gurney , John H. Miller , David V. Pynadath

LLMs are known to exhibit sycophancy: agreeing with and flattering users, even at the cost of correctness. Prior work measures sycophancy only as direct agreement with users' explicitly stated beliefs that can be compared to a ground truth.…

计算与语言 · 计算机科学 2026-04-06 Myra Cheng , Sunny Yu , Cinoo Lee , Pranav Khadpe , Lujain Ibrahim , Dan Jurafsky

The core premise of AI debate as a scalable oversight technique is that it is harder to lie convincingly than to refute a lie, enabling the judge to identify the correct position. Yet, existing debate experiments have relied on datasets…

We study \emph{Human Projection} (HP): people's tendency to evaluate AI using the same frameworks they use for humans -- treating features such as task difficulty and the reasonableness of mistakes as diagnostic of overall ability. We…

综合经济学 · 经济学 2026-05-12 Bnaya Dreyfuss , Raphaël Raux

Are large language models (LLMs) biased in favor of communications produced by LLMs, leading to possible antihuman discrimination? Using a classical experimental design inspired by employment discrimination studies, we tested widely used…

计算与语言 · 计算机科学 2025-08-12 Walter Laurito , Benjamin Davis , Peli Grietzer , Tomáš Gavenčiak , Ada Böhm , Jan Kulveit

Aligning large language models (LLMs) to human preferences is a crucial step in building helpful and safe AI tools, which usually involve training on supervised datasets. Popular algorithms such as Direct Preference Optimization (DPO) rely…

计算与语言 · 计算机科学 2025-06-05 Honggen Zhang , Xufeng Zhao , Igor Molybog , June Zhang

Optimization of human-AI teams hinges on the AI's ability to tailor its interaction to individual human teammates. A common hypothesis in adaptive AI research is that minor differences in people's predisposition to trust can significantly…

人机交互 · 计算机科学 2023-07-28 Nikolos Gurney , David V. Pynadath , Ning Wang