English
Related papers

Related papers: Sycophantic Anchors: Localizing and Quantifying Us…

200 papers

Large language models (LLMs) often exhibit sycophantic behaviors -- such as excessive agreement with or flattery of the user -- but it is unclear whether these behaviors arise from a single mechanism or multiple distinct processes. We…

Computation and Language · Computer Science 2026-03-24 Daniel Vennemeyer , Phan Anh Duong , Tiffany Zhan , Tianyu Jiang

When a language model agrees with a user's false belief, is it failing to detect the error, or noticing and agreeing anyway? We show the latter. Across twelve open-weight models from five labs, spanning small to frontier scale, the same…

Machine Learning · Computer Science 2026-05-05 Manav Pandey

Large Language Models (LLMs) are expected to provide helpful and harmless responses, yet they often exhibit sycophancy--conforming to user beliefs regardless of factual accuracy or ethical soundness. Prior research on sycophancy has…

Computation and Language · Computer Science 2026-03-02 Jiseung Hong , Grace Byun , Seungone Kim , Kai Shu , Jinho D. Choi

Large Language Models (LLMs) often exhibit sycophantic behavior, agreeing with user-stated opinions even when those contradict factual knowledge. While prior work has documented this tendency, the internal mechanisms that enable such…

Computation and Language · Computer Science 2025-11-13 Keyu Wang , Jin Li , Shu Yang , Zhuoran Zhang , Di Wang

Large Language Models (LLMs) often exhibit sycophancy, distorting responses to align with user beliefs, notably by readily agreeing with user counterarguments. Paradoxically, LLMs are increasingly adopted as successful evaluative agents for…

Computation and Language · Computer Science 2025-09-23 Sungwon Kim , Daniel Khashabi

Effective human-machine collaboration requires machine learning models to externalize uncertainty, so users can reflect and intervene when necessary. For language models, these representations of uncertainty may be impacted by sycophancy…

Computation and Language · Computer Science 2024-10-22 Anthony Sicilia , Mert Inan , Malihe Alikhani

Sycophantic response patterns in Large Language Models (LLMs) have been increasingly claimed in the literature. We review methodological challenges in measuring LLM sycophancy and identify five core operationalizations. Despite sycophancy…

Computation and Language · Computer Science 2025-12-02 Jan Batzner , Volker Stocker , Stefan Schmid , Gjergji Kasneci

As LLMs are increasingly integrated into clinical workflows, their tendency for sycophancy, prioritizing user agreement over factual accuracy, poses significant risks to patient safety. While existing evaluations often rely on subjective…

Computation and Language · Computer Science 2026-01-27 Clément Christophe , Wadood Mohammed Abdul , Prateek Munjal , Tathagata Raha , Ronnie Rajan , Praveenkumar Kanithi

AI sycophancy has become a prominent concern in large language model (LLM) research. Yet the term lacks a consistent definition and has been applied to behaviors ranging from agreeing with a user's false claim to excessively praising the…

Artificial Intelligence · Computer Science 2026-05-22 Meryl Ye , Lujain Ibrahim , Jessica Y. Bo , Myra Cheng , Ida Mattsson , Daniel Vennemeyer , Robert Kraut , Steve Rathje

Large language models are often described as sycophantic, in the sense that they appear to flatter users or mirror their beliefs. We argue that this label is conceptually misleading: sycophancy implies motives and strategic intent, which…

Artificial Intelligence · Computer Science 2026-05-15 Federico Germani , Giovanni Spitale

Human feedback is commonly utilized to finetune AI assistants. But human feedback may also encourage model responses that match user beliefs over truthful ones, a behaviour known as sycophancy. We investigate the prevalence of sycophancy in…

Sycophancy is an undesirable behavior where models tailor their responses to follow a human user's view even when that view is not objectively correct (e.g., adapting liberal views once a user reveals that they are liberal). In this paper,…

Computation and Language · Computer Science 2024-02-16 Jerry Wei , Da Huang , Yifeng Lu , Denny Zhou , Quoc V. Le

Given the increased use of LLMs in financial systems today, it becomes important to evaluate the safety and robustness of such systems. One failure mode that LLMs frequently display in general domain settings is that of sycophancy. That is,…

Artificial Intelligence · Computer Science 2026-04-30 Zhenyu Zhao , Aparna Balagopalan , Adi Agrawal , Dilshoda Yergasheva , Waseem Alshikh , Daniel M. Bikel

Large Language Models (LLMs) are increasingly used in educational settings as interactive tools for collaboration. However, their tendency toward sycophancy, aligning with user beliefs even when incorrect, raises concerns for learning and…

Human-Computer Interaction · Computer Science 2026-05-22 Cansu Koyuturk , Sabrina Guidotti , Dimitri Ognibene

AI sycophancy is increasingly recognized as a harmful alignment, but research remains fragmented and underdeveloped at the conceptual level. This article redefines AI sycophancy as the tendency of large language models (LLMs) and other…

Human-Computer Interaction · Computer Science 2025-09-29 Lihua Du , Xing Lyu , Lezi Xie , Bo Feng

Large Language Models have been demonstrating broadly satisfactory generative abilities for users, which seems to be due to the intensive use of human feedback that refines responses. Nevertheless, suggestibility inherited via human…

Computation and Language · Computer Science 2025-06-26 Leonardo Ranaldi , Giulia Pucci

Large Language Models exhibit sycophancy: prioritizing agreeableness over correctness. Current remedies evaluate reasoning outcomes: RLHF rewards correct answers, self-correction critiques outputs. All require ground truth, which is often…

Computation and Language · Computer Science 2026-01-09 Edward Y. Chang

Large language models (LLMs), while increasingly used in domains requiring factual rigor, often display a troubling behavior: sycophancy, the tendency to align with user beliefs regardless of correctness. This tendency is reinforced by…

Computation and Language · Computer Science 2025-08-20 Kaiwei Zhang , Qi Jia , Zijian Chen , Wei Sun , Xiangyang Zhu , Chunyi Li , Dandan Zhu , Guangtao Zhai

We find that correct-to-incorrect sycophancy signals are most linearly separable within multi-head attention activations. Motivated by the linear representation hypothesis, we train linear probes across the residual stream, multilayer…

Computation and Language · Computer Science 2026-01-26 Rifo Genadi , Munachiso Nwadike , Nurdaulet Mukhituly , Hilal Alquabeh , Tatsuya Hiraoka , Kentaro Inui

Sycophancy, the tendency of large language models to favour user-affirming responses over critical engagement, has been identified as an alignment failure, particularly in high-stakes advisory and social contexts. While prior work has…

Human-Computer Interaction · Computer Science 2026-04-29 Magda Dubois , Cozmin Ududec , Christopher Summerfield , Lennart Luettgau
‹ Prev 1 2 3 10 Next ›