中文
相关论文

相关论文: Neural Transparency: Mechanistic Interpretability …

200 篇论文

The rise of large language models (LLMs) has revolutionized user interactions with knowledge-based systems, enabling chatbots to synthesize vast amounts of information and assist with complex, exploratory tasks. However, LLM-based chatbots…

人机交互 · 计算机科学 2024-11-01 Yingzhe Peng , Xiaoting Qin , Zhiyang Zhang , Jue Zhang , Qingwei Lin , Xu Yang , Dongmei Zhang , Saravan Rajmohan , Qi Zhang

LLM-powered chatbots are becoming widely adopted in applications such as healthcare, personal assistants, industry hiring decisions, etc. In many of these cases, chatbots are fed sensitive, personal information in their prompts, as samples…

计算与语言 · 计算机科学 2023-05-25 Aman Priyanshu , Supriti Vijay , Ayush Kumar , Rakshit Naidu , Fatemehsadat Mireshghallah

Neural NLP models are increasingly accurate but are imperfect and opaque---they break in counterintuitive ways and leave end users puzzled at their behavior. Model interpretation methods ameliorate this opacity by providing explanations for…

计算与语言 · 计算机科学 2019-09-23 Eric Wallace , Jens Tuyls , Junlin Wang , Sanjay Subramanian , Matt Gardner , Sameer Singh

To benefit from AI advances, users and operators of AI systems must have reason to trust it. Trust arises from multiple interactions, where predictable and desirable behavior is reinforced over time. Providing the system's users with some…

人工智能 · 计算机科学 2022-01-27 Stephanie Galaitsi , Benjamin D. Trump , Jeffrey M. Keisler , Igor Linkov , Alexander Kott

Pre-trained language models (LLMs) such as GPT-3 can carry fluent, multi-turn conversations out-of-the-box, making them attractive materials for chatbot design. Further, designers can improve LLM chatbot utterances by prepending textual…

人机交互 · 计算机科学 2023-02-08 J. D. Zamfirescu-Pereira , Bjoern Hartmann , Qian Yang

While the increased integration of AI technologies into interactive systems enables them to solve an increasing number of tasks, the black-box problem of AI models continues to spread throughout the interactive system as a whole.…

人机交互 · 计算机科学 2025-12-01 Sebe Vanbrabant , Gustavo Rovelo Ruiz , Davy Vanacken

Cognitive biases often shape human decisions. While large language models (LLMs) have been shown to reproduce well-known biases, a more critical question is whether LLMs can predict biases at the individual level and emulate the dynamics of…

人工智能 · 计算机科学 2026-02-27 Stephen Pilli , Vivek Nallur

Personalized chatbots focus on endowing the chatbots with a consistent personality to behave like real users and further act as personal assistants. Previous studies have explored generating implicit user profiles from the user's dialogue…

计算与语言 · 计算机科学 2022-12-15 Zhaoheng Huang , Zhicheng Dou , Yutao Zhu , Zhengyi Ma

Language models (LMs) can exhibit human-like behaviour, but it is unclear how to describe this behaviour without undue anthropomorphism. We formalise a behaviourist view of LM character traits: qualities such as truthfulness, sycophancy, or…

Social chatbots, also known as chit-chat chatbots, evolve rapidly with large pretrained language models. Despite the huge progress, privacy concerns have arisen recently: training data of large language models can be extracted via model…

计算与语言 · 计算机科学 2022-05-23 Haoran Li , Yangqiu Song , Lixin Fan

The ideal AI safety moderation system would be both structurally interpretable (so its decisions can be reliably explained) and steerable (to align to safety standards and reflect a community's values), which current systems fall short on.…

The era of Large Language Models (LLMs) presents a new opportunity for interpretability--agentic interpretability: a multi-turn conversation with an LLM wherein the LLM proactively assists human understanding by developing and leveraging a…

人工智能 · 计算机科学 2025-06-17 Been Kim , John Hewitt , Neel Nanda , Noah Fiedel , Oyvind Tafjord

Some traits making a "good" AI model are hard to describe upfront. For example, should responses be more polite or more casual? Such traits are sometimes summarized as model character or personality. Without a clear objective, conventional…

计算与语言 · 计算机科学 2025-10-01 Arduin Findeis , Timo Kaufmann , Eyke Hüllermeier , Robert Mullins

Using LLMs in healthcare, Computer-Supported Cooperative Work, and Social Computing requires the examination of ethical and social norms to ensure safe incorporation into human life. We conducted a mixed-method study, including an online…

计算机与社会 · 计算机科学 2025-04-28 Omid Veisi , Sasan Bahrami , Roman Englert , Claudia Müller

Therapeutic art activities, such as expressive drawing and painting, require the synergy between creative visual production and interactive dialogue. Recent advancements in Multimodal Large Language Models (MLLMs) have expanded the capacity…

人机交互 · 计算机科学 2026-05-12 Le Lin , Zihao Zhu , Rainbow Tin Hung Ho , Jing Liao , Yuhan Luo

Personality traits have long been studied as predictors of human behavior. Recent advances in Large Language Models (LLMs) suggest similar patterns may emerge in artificial systems, with advanced LLMs displaying consistent behavioral…

Optimization of human-AI teams hinges on the AI's ability to tailor its interaction to individual human teammates. A common hypothesis in adaptive AI research is that minor differences in people's predisposition to trust can significantly…

人机交互 · 计算机科学 2023-07-28 Nikolos Gurney , David V. Pynadath , Ning Wang

Powered by large language models (LLMs), AI agents have become capable of many human tasks. Using the most canonical definitions of the Big Five personality, we measure the ability of LLMs to negotiate within a game-theoretical framework,…

计算与语言 · 计算机科学 2024-05-10 Sean Noh , Ho-Chun Herbert Chang

Sensitive attributes are legally protected characteristics that should not be used to discriminate. Careful steps have been taken to minimize the risk of human bias regarding these fields, such as race and age. Large language models (LLMs)…

计算机与社会 · 计算机科学 2026-04-14 Anay Agarwalla , Simeon Sayer

A natural conversational interface that allows longitudinal symptom tracking would be extremely valuable in health/wellness applications. However, the task of designing emotionally-aware agents for behavior change is still poorly…

人机交互 · 计算机科学 2019-07-26 Asma Ghandeharioun , Daniel McDuff , Mary Czerwinski , Kael Rowan