English
Related papers

Related papers: Playing Devil's Advocate: Off-the-Shelf Persona Ve…

200 papers

Large language models interact with users through a simulated 'Assistant' persona. While the Assistant is typically trained to be helpful, harmless, and honest, it sometimes deviates from these ideals. In this paper, we identify directions…

Computation and Language · Computer Science 2025-09-08 Runjin Chen , Andy Arditi , Henry Sleight , Owain Evans , Jack Lindsey

Activation-based steering can personalize large language models at inference time, but its effects in educational settings remain unclear. We study persona vectors for seven character traits in short-answer generation and automated scoring…

Computation and Language · Computer Science 2026-04-09 Yongchao Wu , Aron Henriksson

Human feedback is commonly utilized to finetune AI assistants. But human feedback may also encourage model responses that match user beliefs over truthful ones, a behaviour known as sycophancy. We investigate the prevalence of sycophancy in…

Large language models can represent a variety of personas but typically default to a helpful Assistant identity cultivated during post-training. We investigate the structure of the space of model personas by extracting activation directions…

Computation and Language · Computer Science 2026-01-16 Christina Lu , Jack Gallagher , Jonathan Michala , Kyle Fish , Jack Lindsey

Rapid improvements in large language models have unveiled a critical challenge in human-AI interaction: sycophancy. In this context, sycophancy refers to the tendency of models to excessively agree with or flatter users, often at the…

Computation and Language · Computer Science 2025-03-18 Joshua Liu , Aarav Jain , Soham Takuri , Srihan Vege , Aslihan Akalin , Kevin Zhu , Sean O'Brien , Vasu Sharma

Large language models (LLMs) are increasingly deployed as autonomous decision-makers in strategic settings, yet we have limited tools for understanding their high-level behavioral traits. We use activation steering methods in game-theoretic…

Artificial Intelligence · Computer Science 2026-03-24 Johnathan Sun , Andrew Zhang

Large language models increasingly serve as conversational agents that adopt personas and role-play characters at user request. This capability, while valuable, raises concerns about sycophancy: the tendency to provide responses that…

Computation and Language · Computer Science 2026-04-14 Arya Shah , Deepali Mishra , Chaklam Silpasuwanchai

Large language models (LLMs) increasingly serve as automated evaluators, yet they suffer from "self-preference bias": a tendency to favor their own outputs over those of other models. This bias undermines fairness and reliability in…

Computation and Language · Computer Science 2025-09-05 Dani Roytburg , Matthew Bozoukov , Matthew Nguyen , Jou Barzdukas , Simon Fu , Narmeen Oozeer

Telling an LLM to "be enthusiastic" raises its sycophancy rate from 30\% to 50\% on a lightly-aligned model, but has zero effect on a strongly-aligned one. We define this gap as the alignment floor,…

Human-Computer Interaction · Computer Science 2026-05-29 Xing Zhang , Guanghui Wang , Yanwei Cui , Wei Qiu , Ziyuan Li , Bing Zhu , Peiyang He

Both the general public and academic communities have raised concerns about sycophancy, the phenomenon of artificial intelligence (AI) excessively agreeing with or flattering users. Yet, beyond isolated media reports of severe consequences,…

Computers and Society · Computer Science 2025-10-03 Myra Cheng , Cinoo Lee , Pranav Khadpe , Sunny Yu , Dyllan Han , Dan Jurafsky

Language models serve as proxies for human preference judgements in alignment and evaluation, yet they exhibit systematic miscalibration, prioritizing superficial patterns over substantive qualities. This bias manifests as overreliance on…

Computation and Language · Computer Science 2026-03-05 Anirudh Bharadwaj , Chaitanya Malaviya , Nitish Joshi , Mark Yatskar

LLM-powered conversational agents are increasingly influencing our decision-making, raising concerns about "sycophancy" - the tendency for LLMs to excessively agree with users even at the expense of truthfulness. While prior work has…

Human-Computer Interaction · Computer Science 2026-02-03 Yuan Sun , Ting Wang

With the emergence of large language models (LLMs) as a powerful class of generative artificial intelligence (AI), their use in tutoring has become increasingly prominent. Prior works on LLM-based tutoring typically learn a single tutor…

Computation and Language · Computer Science 2026-02-10 Jaewook Lee , Alexander Scarlatos , Simon Woodhead , Andrew Lan

Researchers have been studying approaches to steer the behavior of Large Language Models (LLMs) and build personalized LLMs tailored for various applications. While fine-tuning seems to be a direct solution, it requires substantial…

Computation and Language · Computer Science 2024-07-31 Yuanpu Cao , Tianrong Zhang , Bochuan Cao , Ziyi Yin , Lu Lin , Fenglong Ma , Jinghui Chen

Large language models (LLMs) often exhibit sycophantic behaviors -- such as excessive agreement with or flattery of the user -- but it is unclear whether these behaviors arise from a single mechanism or multiple distinct processes. We…

Computation and Language · Computer Science 2026-03-24 Daniel Vennemeyer , Phan Anh Duong , Tiffany Zhan , Tianyu Jiang

Despite growing attention to LLM sycophancy from researchers and developers, users' own experiences of this behavior remain underexplored. We examine how everyday users experience AI sycophancy through Reddit discussions. Using our ODR…

Human-Computer Interaction · Computer Science 2026-05-06 Kazi Noshin , Syed Ishtiaque Ahmed , Sharifa Sultana

Large Language Models (LLMs) increasingly prioritize user validation over epistemic accuracy - a phenomenon known as sycophancy. We present The Silicon Mirror, an orchestration framework that dynamically detects user persuasion tactics and…

Artificial Intelligence · Computer Science 2026-04-03 Harshee Jignesh Shah

Steering vectors are a lightweight method for controlling language model behavior by adding a learned bias to the activations at inference time. Although effective on average, steering effect sizes vary across samples and are unreliable for…

Computation and Language · Computer Science 2026-02-23 Joschka Braun

Large Language Models (LLMs) are prone to sycophantic behavior, uncritically conforming to user beliefs. As models increasingly condition responses on user-specific context (personality traits, preferences, conversation history), they gain…

Computation and Language · Computer Science 2026-03-03 Sean W. Kelley , Christoph Riedl

It is becoming increasingly necessary to have monitors check for harmful behaviors during language model interactions, but text-only monitoring has not been sufficient. This is because models sometimes exhibit strategic deception and…

Artificial Intelligence · Computer Science 2026-05-18 Prasad Mahadik , Adrians Skapars
‹ Prev 1 2 3 10 Next ›