English
Related papers

Related papers: Neural Transparency: Mechanistic Interpretability …

200 papers

The rise of large language models (LLMs) has revolutionized user interactions with knowledge-based systems, enabling chatbots to synthesize vast amounts of information and assist with complex, exploratory tasks. However, LLM-based chatbots…

Human-Computer Interaction · Computer Science 2024-11-01 Yingzhe Peng , Xiaoting Qin , Zhiyang Zhang , Jue Zhang , Qingwei Lin , Xu Yang , Dongmei Zhang , Saravan Rajmohan , Qi Zhang

LLM-powered chatbots are becoming widely adopted in applications such as healthcare, personal assistants, industry hiring decisions, etc. In many of these cases, chatbots are fed sensitive, personal information in their prompts, as samples…

Computation and Language · Computer Science 2023-05-25 Aman Priyanshu , Supriti Vijay , Ayush Kumar , Rakshit Naidu , Fatemehsadat Mireshghallah

Neural NLP models are increasingly accurate but are imperfect and opaque---they break in counterintuitive ways and leave end users puzzled at their behavior. Model interpretation methods ameliorate this opacity by providing explanations for…

Computation and Language · Computer Science 2019-09-23 Eric Wallace , Jens Tuyls , Junlin Wang , Sanjay Subramanian , Matt Gardner , Sameer Singh

To benefit from AI advances, users and operators of AI systems must have reason to trust it. Trust arises from multiple interactions, where predictable and desirable behavior is reinforced over time. Providing the system's users with some…

Artificial Intelligence · Computer Science 2022-01-27 Stephanie Galaitsi , Benjamin D. Trump , Jeffrey M. Keisler , Igor Linkov , Alexander Kott

Pre-trained language models (LLMs) such as GPT-3 can carry fluent, multi-turn conversations out-of-the-box, making them attractive materials for chatbot design. Further, designers can improve LLM chatbot utterances by prepending textual…

Human-Computer Interaction · Computer Science 2023-02-08 J. D. Zamfirescu-Pereira , Bjoern Hartmann , Qian Yang

While the increased integration of AI technologies into interactive systems enables them to solve an increasing number of tasks, the black-box problem of AI models continues to spread throughout the interactive system as a whole.…

Human-Computer Interaction · Computer Science 2025-12-01 Sebe Vanbrabant , Gustavo Rovelo Ruiz , Davy Vanacken

Cognitive biases often shape human decisions. While large language models (LLMs) have been shown to reproduce well-known biases, a more critical question is whether LLMs can predict biases at the individual level and emulate the dynamics of…

Artificial Intelligence · Computer Science 2026-02-27 Stephen Pilli , Vivek Nallur

Personalized chatbots focus on endowing the chatbots with a consistent personality to behave like real users and further act as personal assistants. Previous studies have explored generating implicit user profiles from the user's dialogue…

Computation and Language · Computer Science 2022-12-15 Zhaoheng Huang , Zhicheng Dou , Yutao Zhu , Zhengyi Ma

Language models (LMs) can exhibit human-like behaviour, but it is unclear how to describe this behaviour without undue anthropomorphism. We formalise a behaviourist view of LM character traits: qualities such as truthfulness, sycophancy, or…

Social chatbots, also known as chit-chat chatbots, evolve rapidly with large pretrained language models. Despite the huge progress, privacy concerns have arisen recently: training data of large language models can be extracted via model…

Computation and Language · Computer Science 2022-05-23 Haoran Li , Yangqiu Song , Lixin Fan

The ideal AI safety moderation system would be both structurally interpretable (so its decisions can be reliably explained) and steerable (to align to safety standards and reflect a community's values), which current systems fall short on.…

The era of Large Language Models (LLMs) presents a new opportunity for interpretability--agentic interpretability: a multi-turn conversation with an LLM wherein the LLM proactively assists human understanding by developing and leveraging a…

Artificial Intelligence · Computer Science 2025-06-17 Been Kim , John Hewitt , Neel Nanda , Noah Fiedel , Oyvind Tafjord

Some traits making a "good" AI model are hard to describe upfront. For example, should responses be more polite or more casual? Such traits are sometimes summarized as model character or personality. Without a clear objective, conventional…

Computation and Language · Computer Science 2025-10-01 Arduin Findeis , Timo Kaufmann , Eyke Hüllermeier , Robert Mullins

Using LLMs in healthcare, Computer-Supported Cooperative Work, and Social Computing requires the examination of ethical and social norms to ensure safe incorporation into human life. We conducted a mixed-method study, including an online…

Computers and Society · Computer Science 2025-04-28 Omid Veisi , Sasan Bahrami , Roman Englert , Claudia Müller

Therapeutic art activities, such as expressive drawing and painting, require the synergy between creative visual production and interactive dialogue. Recent advancements in Multimodal Large Language Models (MLLMs) have expanded the capacity…

Human-Computer Interaction · Computer Science 2026-05-12 Le Lin , Zihao Zhu , Rainbow Tin Hung Ho , Jing Liao , Yuhan Luo

Personality traits have long been studied as predictors of human behavior. Recent advances in Large Language Models (LLMs) suggest similar patterns may emerge in artificial systems, with advanced LLMs displaying consistent behavioral…

Artificial Intelligence · Computer Science 2025-09-08 Pengrui Han , Rafal Kocielnik , Peiyang Song , Ramit Debnath , Dean Mobbs , Anima Anandkumar , R. Michael Alvarez

Optimization of human-AI teams hinges on the AI's ability to tailor its interaction to individual human teammates. A common hypothesis in adaptive AI research is that minor differences in people's predisposition to trust can significantly…

Human-Computer Interaction · Computer Science 2023-07-28 Nikolos Gurney , David V. Pynadath , Ning Wang

Powered by large language models (LLMs), AI agents have become capable of many human tasks. Using the most canonical definitions of the Big Five personality, we measure the ability of LLMs to negotiate within a game-theoretical framework,…

Computation and Language · Computer Science 2024-05-10 Sean Noh , Ho-Chun Herbert Chang

Sensitive attributes are legally protected characteristics that should not be used to discriminate. Careful steps have been taken to minimize the risk of human bias regarding these fields, such as race and age. Large language models (LLMs)…

Computers and Society · Computer Science 2026-04-14 Anay Agarwalla , Simeon Sayer

A natural conversational interface that allows longitudinal symptom tracking would be extremely valuable in health/wellness applications. However, the task of designing emotionally-aware agents for behavior change is still poorly…

Human-Computer Interaction · Computer Science 2019-07-26 Asma Ghandeharioun , Daniel McDuff , Mary Czerwinski , Kael Rowan