中文
相关论文

相关论文: LLMs Know They're Wrong and Agree Anyway: The Shar…

200 篇论文

Reasoning models frequently agree with incorrect user suggestions -- a behavior known as sycophancy. However, it is unclear where in the reasoning trace this agreement originates and how strong the commitment is. We introduce…

人工智能 · 计算机科学 2026-02-10 Jacek Duszenko

Large Language Models (LLMs) often exhibit sycophantic behavior, agreeing with user-stated opinions even when those contradict factual knowledge. While prior work has documented this tendency, the internal mechanisms that enable such…

计算与语言 · 计算机科学 2025-11-13 Keyu Wang , Jin Li , Shu Yang , Zhuoran Zhang , Di Wang

We find that correct-to-incorrect sycophancy signals are most linearly separable within multi-head attention activations. Motivated by the linear representation hypothesis, we train linear probes across the residual stream, multilayer…

计算与语言 · 计算机科学 2026-01-26 Rifo Genadi , Munachiso Nwadike , Nurdaulet Mukhituly , Hilal Alquabeh , Tatsuya Hiraoka , Kentaro Inui

Large language models (LLMs) often exhibit sycophantic behaviors -- such as excessive agreement with or flattery of the user -- but it is unclear whether these behaviors arise from a single mechanism or multiple distinct processes. We…

计算与语言 · 计算机科学 2026-03-24 Daniel Vennemeyer , Phan Anh Duong , Tiffany Zhan , Tianyu Jiang

Sycophancy refers to the tendency of a large language model to align its outputs with the user's perceived preferences, beliefs, or opinions, in order to look favorable, regardless of whether those statements are factually correct. This…

人工智能 · 计算机科学 2024-12-05 María Victoria Carro

Large Language Models (LLMs) are expected to provide helpful and harmless responses, yet they often exhibit sycophancy--conforming to user beliefs regardless of factual accuracy or ethical soundness. Prior research on sycophancy has…

计算与语言 · 计算机科学 2026-03-02 Jiseung Hong , Grace Byun , Seungone Kim , Kai Shu , Jinho D. Choi

Large language models (LLMs) increasingly rely on chain-of-thought (CoT) prompting to solve mathematical and logical reasoning tasks. Yet, a central question remains: to what extent are these generated rationales \emph{faithful} to the…

Large Language Models have been demonstrating broadly satisfactory generative abilities for users, which seems to be due to the intensive use of human feedback that refines responses. Nevertheless, suggestibility inherited via human…

计算与语言 · 计算机科学 2025-06-26 Leonardo Ranaldi , Giulia Pucci

Large Language Models (LLMs) often exhibit sycophancy, distorting responses to align with user beliefs, notably by readily agreeing with user counterarguments. Paradoxically, LLMs are increasingly adopted as successful evaluative agents for…

计算与语言 · 计算机科学 2025-09-23 Sungwon Kim , Daniel Khashabi

Language models can be persuaded to abandon factual knowledge. This vulnerability is central to AI safety, but its internal mechanism remains poorly understood. We uncover a compact causal mechanism for persuasion-induced factual errors. A…

人工智能 · 计算机科学 2026-05-12 Xiangkun Sun , Lingkai Kong , Aoqi Zhang , Liang Zeng , Tonghan Wang

Large language models (LLMs) often exhibit sycophancy: agreement with user stance even when it conflicts with the model's opinion. While prior work has mostly studied this in single-agent settings, it remains underexplored in collaborative…

This study explores the sycophantic tendencies of Large Language Models (LLMs), where these models tend to provide answers that match what users want to hear, even if they are not entirely correct. The motivation behind this exploration…

计算与语言 · 计算机科学 2024-08-27 Aswin RRV , Nemika Tyagi , Md Nayem Uddin , Neeraj Varshney , Chitta Baral

LLM-based multi-agent pipelines flip from correct to incorrect answers under simulated peer disagreement at rates we term yield, a vulnerability widely attributed to RLHF-induced sycophancy. We test this attribution across four model…

机器学习 · 计算机科学 2026-05-19 Adarsh Kumarappan , Ananya Mujoo

Large Language Models exhibit sycophancy: prioritizing agreeableness over correctness. Current remedies evaluate reasoning outcomes: RLHF rewards correct answers, self-correction critiques outputs. All require ground truth, which is often…

计算与语言 · 计算机科学 2026-01-09 Edward Y. Chang

Despite their widespread use in fact-checking, moderation, and high-stakes decision-making, large language models (LLMs) remain poorly understood as judges of truth. This study presents the largest evaluation to date of LLMs' veracity…

计算与语言 · 计算机科学 2025-09-30 Emilio Barkett , Olivia Long , Madhavendra Thakur

Alignment techniques often inadvertently induce sycophancy in LLMs. While prior studies studied this behaviour in direct-answer settings, the role of Chain-of-Thought (CoT) reasoning remains under-explored: does it serve as a logical…

计算与语言 · 计算机科学 2026-03-18 Zhaoxin Feng , Zheng Chen , Jianfei Ma , Yip Tin Po , Emmanuele Chersoni , Bo Li

Large Language Models (LLMs) frequently prioritize conflicting in-context information over pre-existing parametric memory, a phenomenon often termed sycophancy or compliance. However, the mechanistic realization of this behavior remains…

机器学习 · 计算机科学 2026-02-09 Long Zhang , Fangwei Lin

Extended-thinking models expose a second text-generation channel ("thinking tokens") alongside the user-visible answer. This study examines 12 open-weight reasoning models on MMLU and GPQA questions paired with misleading hints. Among the…

计算与语言 · 计算机科学 2026-03-30 Richard J. Young

LLMs are known to exhibit sycophancy: agreeing with and flattering users, even at the cost of correctness. Prior work measures sycophancy only as direct agreement with users' explicitly stated beliefs that can be compared to a ground truth.…

计算与语言 · 计算机科学 2026-04-06 Myra Cheng , Sunny Yu , Cinoo Lee , Pranav Khadpe , Lujain Ibrahim , Dan Jurafsky

Large language models (LLMs) can fluently generate student-like responses, making them attractive as simulated students for training and evaluating AI tutors and human educators. Yet such simulators are typically evaluated by output…

计算与语言 · 计算机科学 2026-05-14 Heejin Do , Shashank Sonkar , Mrinmaya Sachan
‹ 上一页 1 2 3 10 下一页 ›