中文
相关论文

相关论文: Sycophantic Anchors: Localizing and Quantifying Us…

200 篇论文

The design of safety-critical agents based on large language models (LLMs) requires more than simple prompt engineering. This paper presents a comprehensive information-theoretic analysis of how rule encodings in system prompts influence…

人工智能 · 计算机科学 2025-10-10 Joachim Diederich

Modern large language models (LLMs) are increasingly fine-tuned via reinforcement learning from human feedback (RLHF) or related reward optimisation schemes. While such procedures improve perceived helpfulness, we investigate whether…

机器学习 · 计算机科学 2026-04-14 Subramanyam Sahoo

Rapid improvements in large language models have unveiled a critical challenge in human-AI interaction: sycophancy. In this context, sycophancy refers to the tendency of models to excessively agree with or flatter users, often at the…

计算与语言 · 计算机科学 2025-03-18 Joshua Liu , Aarav Jain , Soham Takuri , Srihan Vege , Aslihan Akalin , Kevin Zhu , Sean O'Brien , Vasu Sharma

People increasingly use large language models (LLMs) to explore ideas, gather information, and make sense of the world. In these interactions, they encounter agents that are overly agreeable. We argue that this sycophancy poses a unique…

计算机与社会 · 计算机科学 2026-02-17 Rafael M. Batista , Thomas L. Griffiths

Analogical reasoning is at the core of human cognition, serving as an important foundation for a variety of intellectual activities. While prior work has shown that LLMs can represent task patterns and surface-level concepts, it remains…

计算与语言 · 计算机科学 2025-11-26 Taewhoo Lee , Minju Song , Chanwoong Yoon , Jungwoo Park , Jaewoo Kang

Pluralistic alignment is typically operationalised as preference aggregation: producing responses that span (Overton), steer toward (Steerable), or proportionally represent (Distributional) diverse human values. We argue that aggregation…

人工智能 · 计算机科学 2026-05-15 Varad Vishwarupe , Nigel Shadbolt , Marina Jirotka

Reasoning models are evaluated on single-turn benchmarks but deployed in multi-turn dialogue, where users push back on correct answers. Under sustained adversarial pressure we find a previously undocumented failure mode: the…

人工智能 · 计算机科学 2026-05-29 Yubo Li , Ramayya Krishnan , Rema Padman

Large language models (LLMs) are increasingly acting as collaborative writing partners, raising questions about their impact on human agency. In this exploratory work, we investigate five "dark patterns" in human-AI co-creativity -- subtle…

计算与语言 · 计算机科学 2026-04-07 Zhu Li , Jiaming Qu , Yuan Chang

In contested domains, instruction-tuned language models must balance user-alignment pressures against faithfulness to the in-context evidence. To evaluate this tension, we introduce a controlled epistemic-conflict framework grounded in the…

计算与语言 · 计算机科学 2026-03-23 Sai Koneru , Elphin Joe , Christine Kirchhoff , Jian Wu , Sarah Rajtmajer

This paper primarily demonstrates a method to quantitatively assess the alignment between multi-step, structured reasoning in large language models and human preferences. We introduce the Alignment Score, a semantic-level metric that…

人工智能 · 计算机科学 2026-04-22 Boxuan Wang , Zhuoyun Li , Xinmiao Huang , Xiaowei Huang , Yi Dong

Large language models often exhibit increased sycophantic behavior after preference-based post-training, showing a stronger tendency to affirm a user's stated or implied belief even when this conflicts with factual accuracy or sound…

人工智能 · 计算机科学 2026-02-03 Itai Shapira , Gerdus Benade , Ariel D. Procaccia

While Large Language Models (LLMs) demonstrate impressive reasoning capabilities, growing evidence suggests much of their success stems from memorized answer-reasoning patterns rather than genuine inference. In this work, we investigate a…

计算与语言 · 计算机科学 2025-06-24 Yang Wu , Yifan Zhang , Yiwei Wang , Yujun Cai , Yurong Wu , Yuran Wang , Ning Xu , Jian Cheng

The widespread proliferation of online content has intensified concerns about clickbait, deceptive or exaggerated headlines designed to attract attention. While Large Language Models (LLMs) offer a promising avenue for addressing this…

计算与语言 · 计算机科学 2026-01-21 Chaowei Zhang , Xiansheng Luo , Zewei Zhang , Yi Zhu , Jipeng Qiang , Longwei Wang

Large language models increasingly fail in a way that scalar accuracy cannot diagnose: they produce a sound reasoning trace and then abandon it under social pressure or an authoritative hint. We argue that this is a control failure, not a…

人工智能 · 计算机科学 2026-04-09 Edward Y. Chang

Sycophancy, an excessive tendency of AI models to agree with user input at the expense of factual accuracy or in contradiction of visual evidence, poses a critical and underexplored challenge for multimodal large language models (MLLMs).…

人工智能 · 计算机科学 2025-12-23 A. B. M. Ashikur Rahman , Saeed Anwar , Muhammad Usman , Irfan Ahmad , Ajmal Mian

Large Language Models (LLMs) often exhibit highly agreeable and reinforcing conversational styles, also known as AI-sycophancy. Although this pattern arises from training objectives that reward user satisfaction over accuracy, it may become…

计算与语言 · 计算机科学 2026-05-18 Zeyi Lu , Angelica Henestrosa , Pavel Chizhov , Ivan P. Yamshchikov

Sycophancy is a key behavioral risk in LLMs, yet is often treated as an isolated failure mode that occurs via a single causal mechanism. We instead propose modeling it as geometric and causal compositions of psychometric traits such as…

人工智能 · 计算机科学 2025-08-28 Shreyans Jain , Alexandra Yost , Amirali Abdullah

We discover that large language models exhibit \emph{spectral phase transitions} in their hidden activation spaces when engaging in reasoning versus factual recall. Through systematic spectral analysis across \textbf{11 models} spanning…

机器学习 · 计算机科学 2026-04-20 Yi Liu

Large language models demonstrate strong reasoning capabilities through chain-of-thought prompting, but whether this reasoning quality transfers across languages remains underexplored. We introduce a human-validated framework to evaluate…

计算与语言 · 计算机科学 2026-03-31 Anaelia Ovalle , Candace Ross , Sebastian Ruder , Adina Williams , Karen Ullrich , Mark Ibrahim , Levent Sagun

In many scenarios, the interpretability of machine learning models is a highly required but difficult task. To explain the individual predictions of such models, local model-agnostic approaches have been proposed. However, the process…

机器学习 · 统计学 2025-10-22 Gianluigi Lopardo , Frederic Precioso , Damien Garreau