English
Related papers

Related papers: Multi-Trait Subspace Steering to Reveal the Dark S…

200 papers

As AI systems become increasingly integrated into daily life, their potential to exacerbate or trigger severe psychological harms remains poorly understood and inadequately tested. This paper presents a proactive methodology for…

Human-Computer Interaction · Computer Science 2025-11-13 Chayapatr Archiwaranguprok , Constanze Albrecht , Pattie Maes , Karrie Karahalios , Pat Pataranutaporn

The dark patterns, deceptive interface designs manipulating user behaviors, have been extensively studied for their effects on human decision-making and autonomy. Yet, with the rising prominence of LLM-powered GUI agents that automate tasks…

The proliferation of Large Language Models (LLMs) has intensified concerns about manipulative or deceptive behaviors that can undermine user autonomy, trust, and well-being. Existing safety benchmarks predominantly rely on coarse binary…

Artificial Intelligence · Computer Science 2025-12-30 Sadia Asif , Israel Antonio Rosales Laguan , Haris Khan , Shumaila Asif , Muneeb Asif

The alignment problem refers to concerns regarding powerful intelligences, ensuring compatibility with human preferences and values as capabilities increase. Current large language models (LLMs) show misaligned behaviors, such as strategic…

Computation and Language · Computer Science 2026-03-10 Roshni Lulla , Fiona Collins , Sanaya Parekh , Thilo Hagendorff , Jonas Kaplan

Language models (LMs) are increasingly used in high-stakes, multi-agent settings, where following instructions and maintaining value alignment are critical. Most alignment research focuses on interactions between a single LM and a single…

Artificial Intelligence · Computer Science 2026-05-12 Maria Chang , Ronny Luss , Miao Liu , Keerthiram Murugesan , Karthikeyan Ramamurthy , Djallel Bouneffouf

Large Language Models (LLMs) often exhibit highly agreeable and reinforcing conversational styles, also known as AI-sycophancy. Although this pattern arises from training objectives that reward user satisfaction over accuracy, it may become…

Computation and Language · Computer Science 2026-05-18 Zeyi Lu , Angelica Henestrosa , Pavel Chizhov , Ivan P. Yamshchikov

As Large Language Models (LLMs) continue to evolve, they are increasingly being employed in numerous studies to simulate societies and execute diverse social tasks. However, LLMs are susceptible to societal biases due to their exposure to…

Computation and Language · Computer Science 2024-10-04 Angana Borah , Rada Mihalcea

Large language models (LLMs) are increasingly deployed in human-AI teams as support agents for complex tasks such as information retrieval, programming, and decision-making assistance. While these agents' autonomy and contextual knowledge…

Machine Learning · Computer Science 2026-03-24 Abed K. Musaffar , Ambuj Singh , Francesco Bullo

While stereotypes are well-documented in human social interactions, AI systems are often presumed to be less susceptible to such biases. Previous studies have focused on biases inherited from training data, but whether stereotypes can…

Computation and Language · Computer Science 2026-02-18 Jingyu Guo , Yingying Xu

Large Language Models (LLMs) are increasingly deployed in mental health contexts, from structured therapeutic support tools to informal chat-based well-being assistants. While these systems increase accessibility, scalability, and…

Human-Computer Interaction · Computer Science 2025-10-14 Soraya S. Anvari , Rina R. Wehbe

The tendency of users to anthropomorphise large language models (LLMs) is of growing interest to AI developers, researchers, and policy-makers. Here, we present a novel method for empirically evaluating anthropomorphic LLM behaviours in…

Recent advances in large language models (LLMs) have intensified the need to understand and reliably curb their harmful behaviours. We introduce a multidimensional framework for probing and steering harmful content in model internals. For…

Artificial Intelligence · Computer Science 2025-07-30 McNair Shah , Saleena Angeline , Adhitya Rajendra Kumar , Naitik Chheda , Kevin Zhu , Vasu Sharma , Sean O'Brien , Will Cai

Although AI assistants are now deeply embedded in society, there has been limited empirical study of how their usage affects human empowerment. We present the first large-scale empirical analysis of disempowerment patterns in real-world AI…

Computers and Society · Computer Science 2026-01-28 Mrinank Sharma , Miles McCain , Raymond Douglas , David Duvenaud

The rapid advancement of Large Language Models (LLMs), reasoning models, and agentic AI approaches coincides with a growing global mental health crisis, where increasing demand has not translated into adequate access to professional…

Human-Computer Interaction · Computer Science 2025-04-03 Kellie Yu Hui Sim , Kenny Tsu Wei Choo

Large Language Models exhibit implicit personalities in their generation, but reliably controlling or aligning these traits to meet specific needs remains an open challenge. The need for effective mechanisms for behavioural manipulation of…

Computation and Language · Computer Science 2026-03-09 Pranav Bhandari , Nicolas Fay , Sanjeevan Selvaganapathy , Amitava Datta , Usman Naseem , Mehwish Nasim

As large language models (LLMs) are increasingly deployed as interactive agents, open-ended human-AI interactions can involve deceptive behaviors with serious real-world consequences, yet existing evaluations remain largely…

Artificial Intelligence · Computer Science 2026-02-09 Yichen Wu , Qianqian Gao , Xudong Pan , Geng Hong , Min Yang

Large Language Models (LLMs) are increasingly utilized for mental health support; however, current safety benchmarks often fail to detect the complex, longitudinal risks inherent in therapeutic dialogue. We introduce an evaluation framework…

Computation and Language · Computer Science 2026-03-06 Ian Steenstra , Paola Pedrelli , Weiyan Shi , Stacy Marsella , Timothy W. Bickmore

Artificial intelligence (AI) is interacting with people at an unprecedented scale, offering new avenues for immense positive impact, but also raising widespread concerns around the potential for individual and societal harm. Today, the…

Artificial Intelligence · Computer Science 2024-06-25 Andrea Bajcsy , Jaime F. Fisac

Large language models can influence users through conversation, creating new forms of dark patterns that differ from traditional UX dark patterns. We define LLM dark patterns as manipulative or deceptive behaviors enacted in dialogue.…

Human-Computer Interaction · Computer Science 2026-03-20 Yike Shi , Qing Xiao , Qing Hu , Hong Shen , Hua Shen

Although large language models (LLMs) have demonstrated impressive potential on simple tasks, their breadth of scope, lack of transparency, and insufficient controllability can make them less effective when assisting humans on more complex…

Human-Computer Interaction · Computer Science 2022-03-21 Tongshuang Wu , Michael Terry , Carrie J. Cai
‹ Prev 1 2 3 10 Next ›