English
Related papers

Related papers: Towards Understanding Sycophancy in Language Model…

200 papers

Interfaces for interacting with large language models (LLMs) are often designed to mimic human conversations, typically presenting a single response to user queries. This design choice can obscure the probabilistic and predictive nature of…

Human-Computer Interaction · Computer Science 2025-03-21 Chelse Swoopes , Tyler Holloway , Elena L. Glassman

AI systems increasingly assist human decision making by producing preliminary assessments of complex inputs. However, such AI-generated assessments can often be noisy or systematically biased, raising a central question: how should costly…

Machine Learning · Statistics 2026-03-17 Lezhi Tan , Naomi Sagan , Lihua Lei , Jose Blanchet

As AI becomes more deeply embedded in knowledge work, building assistants that support human creativity and expertise becomes more important. Yet achieving synergy in human-AI collaboration is not easy. Providing AI with detailed…

Human-Computer Interaction · Computer Science 2026-04-22 Sean Kelley , David De Cremer , Christoph Riedl

The conformity effect describes the tendency of individuals to align their responses with the majority. Studying this bias in large language models (LLMs) is crucial, as LLMs are increasingly used in various information-seeking and…

Computation and Language · Computer Science 2025-05-27 Xiaochen Zhu , Caiqi Zhang , Tom Stafford , Nigel Collier , Andreas Vlachos

We fine-tune large language models to write natural language critiques (natural language critical comments) using behavioral cloning. On a topic-based summarization task, critiques written by our models help humans find flaws in summaries…

Computation and Language · Computer Science 2022-06-15 William Saunders , Catherine Yeh , Jeff Wu , Steven Bills , Long Ouyang , Jonathan Ward , Jan Leike

This study examines how user-provided suggestions affect Large Language Models (LLMs) in a simulated educational context, where sycophancy poses significant risks. Testing five different LLMs from the OpenAI GPT-4o and GPT-4.1 model classes…

Computation and Language · Computer Science 2025-06-13 Chuck Arvin

AI assistance in decision-making has become popular, yet people's inappropriate reliance on AI often leads to unsatisfactory human-AI collaboration performance. In this paper, through three pre-registered, randomized human subject…

Human-Computer Interaction · Computer Science 2024-01-17 Zhuoran Lu , Dakuo Wang , Ming Yin

Polite speech poses a fundamental alignment challenge for large language models (LLMs). Humans deploy a rich repertoire of linguistic strategies to balance informational and social goals -- from positive approaches that build rapport…

Computation and Language · Computer Science 2025-10-31 Haoran Zhao , Robert D. Hawkins

Adjusting robot behavior to human preferences can require intensive human feedback, preventing quick adaptation to new users and changing circumstances. Moreover, current approaches typically treat user preferences as a reward, which…

Robotics · Computer Science 2024-10-21 Jakob Thumm , Christopher Agia , Marco Pavone , Matthias Althoff

One way to personalize and steer generations from large language models (LLM) is to assign a persona: a role that describes how the user expects the LLM to behave (e.g., a helpful assistant, a teacher, a woman). This paper investigates how…

Computation and Language · Computer Science 2025-07-02 Pedro Henrique Luz de Araujo , Benjamin Roth

Human-AI collaboration for decision-making strives to achieve team performance that exceeds the performance of humans or AI alone. However, many factors can impact success of Human-AI teams, including a user's domain expertise, mental…

AI systems are fallible, and humans can make mistakes in deciding whether to trust AI over their own judgment. Thus, improving human-AI collaboration requires understanding when, why, and how humans decide to rely on AI. We study two…

Artificial Intelligence · Computer Science 2026-05-28 Maharshi Gor , Yoo Yeon Sung , Yu Hou , Eve Fleisig , Irene Ying , Tianyi Zhou , Jordan Boyd-Graber

Many hyper-personalized AI systems profile people's characteristics (e.g., personality traits) to provide personalized recommendations. These systems are increasingly used to facilitate interactions among people, such as providing teammate…

Human-Computer Interaction · Computer Science 2024-05-28 Qiaosi Wang , Chidimma L. Anyi , Vedant Das Swain , Ashok K. Goel

Aligning large language models with human preferences is critical for creating reliable and controllable AI systems. A human preference can be visualized as a high-dimensional vector where different directions represent trade-offs between…

Computation and Language · Computer Science 2026-02-26 Ruochen Mao , Yuling Shi , Xiaodong Gu , Jiaheng Wei

LLMs are aligned to follow input instructions by learning which of two responses users prefer for a prompt. However, such preference data do not convey why users prefer responses that are chosen or rejected, so LLMs trained on these…

Computation and Language · Computer Science 2025-06-03 Nishant Balepur , Vishakh Padmakumar , Fumeng Yang , Shi Feng , Rachel Rudinger , Jordan Lee Boyd-Graber

Despite their widespread use in fact-checking, moderation, and high-stakes decision-making, large language models (LLMs) remain poorly understood as judges of truth. This study presents the largest evaluation to date of LLMs' veracity…

Computation and Language · Computer Science 2025-09-30 Emilio Barkett , Olivia Long , Madhavendra Thakur

Deep generative models have shown impressive results in text-to-image synthesis. However, current text-to-image models often generate images that are inadequately aligned with text prompts. We propose a fine-tuning method for aligning such…

AI agents designed to collaborate with people benefit from models that enable them to anticipate human behavior. However, realistic models tend to require vast amounts of human data, which is often hard to collect. A good prior or…

Machine Learning · Computer Science 2022-11-22 Mesut Yang , Micah Carroll , Anca Dragan

Large Language Models (LLMs) tend to prioritize adherence to user prompts over providing veracious responses, leading to the sycophancy issue. When challenged by users, LLMs tend to admit mistakes and provide inaccurate responses even if…

Computation and Language · Computer Science 2025-02-06 Wei Chen , Zhen Huang , Liang Xie , Binbin Lin , Houqiang Li , Le Lu , Xinmei Tian , Deng Cai , Yonggang Zhang , Wenxiao Wang , Xu Shen , Jieping Ye

Large language models are often ranked according to their level of alignment with human preferences -- a model is better than other models if its outputs are more frequently preferred by humans. One of the popular ways to elicit human…

Machine Learning · Computer Science 2024-12-05 Ivi Chatzi , Eleni Straitouri , Suhas Thejaswi , Manuel Gomez Rodriguez
‹ Prev 1 8 9 10 Next ›