中文
相关论文

相关论文: Neural Transparency: Mechanistic Interpretability …

200 篇论文

Large Language Models (LLMs) are increasingly being utilized by both candidates and employers in the recruitment context. However, with this comes numerous ethical concerns, particularly related to the lack of transparency in these…

计算与语言 · 计算机科学 2024-02-16 Airlie Hilliard , Cristian Munoz , Zekun Wu , Adriano Soares Koshiyama

Deep neural networks form the backbone of artificial intelligence research, with potential to transform the human experience in areas ranging from autonomous driving to personal assistants, healthcare to education. However, their…

机器学习 · 计算机科学 2025-05-29 Vinitra Swamy

A central goal of explainable artificial intelligence (XAI) is to improve the trust relationship in human-AI interaction. One assumption underlying research in transparent AI systems is that explanations help to better assess predictions of…

人工智能 · 计算机科学 2021-06-23 Felix Biessmann , Viktor Treu

Large Language Models (LLMs) excel at producing broadly relevant text, but this generality becomes a limitation when user-specific preferences are required, such as recommending restaurants or planning travel. In these scenarios, users…

Service and assistive robots are increasingly being deployed in dynamic social environments; however, ensuring transparent and explainable interactions remains a significant challenge. This paper presents a multimodal explainability module…

机器人学 · 计算机科学 2026-04-09 Oluwadamilola Sotomi , Devika Kodi , Aliasghar Arab

Large language models (LLMs) are currently at the forefront of intertwining AI systems with human communication and everyday life. Therefore, it is of great importance to evaluate their emerging abilities. In this study, we show that LLMs,…

计算与语言 · 计算机科学 2023-10-10 Thilo Hagendorff , Sarah Fabi

Large language models, pivotal in artificial intelligence, find diverse applications. ChatGPT (Chat Generative Pre-trained Transformer), an OpenAI creation, stands out as a widely adopted, powerful tool. It excels in chatbots, content…

计算与语言 · 计算机科学 2025-05-28 Walid Hariri

Social chatbots based on large language models are increasingly embedded in everyday platforms, yet how users develop trust in these systems over time remains unclear. We present a four-week longitudinal qualitative survey study (N = 27) of…

计算机与社会 · 计算机科学 2026-04-27 Annie Landerberg , Kari Flatmo , Alan Said

Personalized chatbots focus on endowing chatbots with a consistent personality to behave like real users, give more informative responses, and further act as personal assistants. Existing personalized approaches tried to incorporate several…

计算与语言 · 计算机科学 2021-09-03 Zhengyi Ma , Zhicheng Dou , Yutao Zhu , Hanxun Zhong , Ji-Rong Wen

Transparency and security are both central to Responsible AI, but they may conflict in adversarial settings. We investigate the strategic effect of transparency for agents through the lens of transferable adversarial example attacks. In…

机器学习 · 计算机科学 2025-11-18 Lucas Fenaux , Christopher Srinivasa , Florian Kerschbaum

Our everyday interactions with pervasive systems generate traces that capture various aspects of human behavior and enable machine learning algorithms to extract latent information about users. In this paper, we propose a machine learning…

机器学习 · 统计学 2019-06-06 Benjamin Baron , Mirco Musolesi

AI agents negotiate and transact in natural language with unfamiliar counterparts: a buyer bot facing an unknown seller, or a procurement assistant negotiating with a supplier. In such interactions, the counterpart's LLM, prompts, control…

机器学习 · 计算机科学 2026-05-13 Eilam Shapira , Moshe Tennenholtz , Roi Reichart

Transformer language models are state of the art in a multitude of NLP tasks. Despite these successes, their opaqueness remains problematic. Recent methods aiming to provide interpretability and explainability to black-box models primarily…

计算与语言 · 计算机科学 2022-03-14 Felix Friedrich , Patrick Schramowski , Christopher Tauchmann , Kristian Kersting

Personalized support is essential to fulfill individuals' emotional needs and sustain their mental well-being. Large language models (LLMs), with great customization flexibility, hold promises to enable individuals to create their own…

人机交互 · 计算机科学 2025-05-01 Xi Zheng , Zhuoyang Li , Xinning Gui , Yuhan Luo

Training AI models is challenging, particularly when crafting behavior instructions. Traditional methods rely on machines (supervised learning) or manual pattern discovery, which results in not interpretable models or time sink. While Large…

人机交互 · 计算机科学 2025-03-07 Soya Park , J. D. Zamfirescu-Pereira , Chinmay Kulkarni

Existing approaches for the design of interpretable agent behavior consider different measures of interpretability in isolation. In this paper we posit that, in the design and deployment of human-aware agents in the real world, notions of…

AI agents are promising for high-stakes enterprise workflows, but dependable deployment remains limited because tool-use failures are difficult to diagnose and control. Agents may skip required tool calls, invoke tools unnecessarily, or…

人工智能 · 计算机科学 2026-05-22 Hariom Tatsat , Ariye Shater

Individuals are turning to increasingly anthropomorphic, general-purpose chatbots for AI companionship, rather than roleplay-specific platforms. However, not much is known about how individuals perceive and conduct their relationships with…

Natural language processing (NLP) models often replicate or amplify social bias from training data, raising concerns about fairness. At the same time, their black-box nature makes it difficult for users to recognize biased predictions and…

计算与语言 · 计算机科学 2026-02-12 Yifan Wang , Mayank Jobanputra , Ji-Ung Lee , Soyoung Oh , Isabel Valera , Vera Demberg

As artificial intelligence rapidly transforms society, developers and policymakers struggle to anticipate which applications will face public moral resistance. We propose that these judgments are not idiosyncratic but systematic and…

计算机与社会 · 计算机科学 2025-10-08 Kimmo Eriksson , Simon Karlsson , Irina Vartanova , Pontus Strimling