中文
相关论文

相关论文: Minimum Levels of Interpretability for Artificial …

200 篇论文

Artificial Intelligence (AI) has continued to achieve tremendous success in recent times. However, the decision logic of these frameworks is often not transparent, making it difficult for stakeholders to understand, interpret or explain…

机器学习 · 计算机科学 2025-01-20 Fuseini Mumuni , Alhassan Mumuni

Mechanistic interpretability (MI) is an emerging framework for interpreting neural networks. Given a task and model, MI aims to discover a succinct algorithmic process, an interpretation, that explains the model's decision process on that…

机器学习 · 计算机科学 2026-04-01 Alan Sun , Mariya Toneva

How to attribute responsibility for autonomous artificial intelligence (AI) systems' actions has been widely debated across the humanities and social science disciplines. This work presents two experiments ($N$=200 each) that measure…

计算机与社会 · 计算机科学 2021-02-02 Gabriel Lima , Nina Grgić-Hlača , Meeyoung Cha

Automated interpretability systems aim to reduce the need for human labor and scale analysis to increasingly large models and diverse tasks. Recent efforts toward this goal leverage large language models (LLMs) at increasing levels of…

The problem of human trust in artificial intelligence is one of the most fundamental problems in applied machine learning. Our processes for evaluating AI trustworthiness have substantial ramifications for ML's impact on science, health,…

机器学习 · 计算机科学 2022-02-14 Max W. Shen

Machine learning is an important tool for decision making, but its ethical and responsible application requires rigorous vetting of its interpretability and utility: an understudied problem, particularly for natural language processing…

人工智能 · 计算机科学 2019-06-11 Shi Feng , Jordan Boyd-Graber

The increasing demand in artificial intelligence (AI) for models that are both effective and explainable is critical in domains where safety and trust are paramount. In this study, we introduce MIRA, a transparent and interpretable…

人工智能 · 计算机科学 2024-10-18 Sarah Seifi , Tobias Sukianto , Cecilia Carbonelli , Lorenzo Servadei , Robert Wille

Responsible AI demands systems whose behavioral tendencies can be effectively measured, audited, and adjusted to prevent inadvertently nudging users toward risky decisions or embedding hidden biases in risk aversion. As language models…

Existing approaches for generating human-aware agent behaviors have considered different measures of interpretability in isolation. Further, these measures have been studied under differing assumptions, thus precluding the possibility of…

人工智能 · 计算机科学 2021-04-23 Sarath Sreedharan , Anagha Kulkarni , David E. Smith , Subbarao Kambhampati

As the complexity of AI systems and their interactions with the world increases, generating explanations for their behaviour is important for safely deploying AI. For agents, the most natural abstractions for predicting behaviour attribute…

人工智能 · 计算机科学 2025-06-05 Alexis Bellot , Jonathan Richens , Tom Everitt

In this paper we describe moral quasi-dilemmas (MQDs): situations similar to moral dilemmas, but in which an agent is unsure whether exploring the plan space or the world may reveal a course of action that satisfies all moral requirements.…

人工智能 · 计算机科学 2018-07-10 Daniel Kasenberg , Vasanth Sarathy , Thomas Arnold , Matthias Scheutz , Tom Williams

As AI agents built on large language models (LLMs) become increasingly embedded in society, issues of coordination, control, delegation, and accountability are entangled with concerns over their reliability. To design and implement LLM…

计算机与社会 · 计算机科学 2025-12-09 R. Patrick Xian , Garry A. Gabison , Ahmed Alaa , Christoph Riedl , Grigorios G. Chrysos

Explainability in AI and ML models is critical for fostering trust, ensuring accountability, and enabling informed decision making in high stakes domains. Yet this objective is often unmet in practice. This paper proposes a general purpose…

统计金融 · 定量金融 2025-09-03 N. Jean , G. Le Pera

Artificial Intelligence (AI) has a tremendous impact on the unexpected growth of technology in almost every aspect. AI-powered systems are monitoring and deciding about sensitive economic and societal issues. The future is towards…

机器学习 · 计算机科学 2022-06-14 Ioannis Mollas , Nick Bassiliades , Grigorios Tsoumakas

This article proposes a new integration of linguistic anthropology and machine learning (ML) around convergent interests in both the underpinnings of language and making language technologies more socially responsible. While linguistic…

计算机与社会 · 计算机科学 2024-11-11 Graham M. Jones , Shai Satran , Arvind Satyanarayan

Is it possible to evaluate the moral cognition of complex artificial agents? In this work, we take a look at one aspect of morality: `doing the right thing for the right reasons.' We propose a behavior-based analysis of artificial moral…

As artificial intelligence systems increasingly inform high-stakes decisions across sectors, transparency has become foundational to responsible and trustworthy AI implementation. Leveraging our role as a leading institute in advancing AI…

机器学习 · 计算机科学 2025-08-01 Dhanesh Ramachandram , Himanshu Joshi , Judy Zhu , Dhari Gandhi , Lucas Hartman , Ananya Raval

We consider two fundamental and related issues currently faced by Artificial Intelligence (AI) development: the lack of ethics and interpretability of AI decisions. Can interpretable AI decisions help to address ethics in AI? Using a…

人工智能 · 计算机科学 2021-09-21 Jean-Marie John-Mathews

The concepts of blameworthiness and wrongness are of fundamental importance in human moral life. But to what extent are humans disposed to blame artificially intelligent agents, and to what extent will they judge their actions to be morally…

计算机与社会 · 计算机科学 2021-02-09 Michael T. Stuart , Markus Kneer

AI practitioners increasingly use large language model (LLM) agents in compound AI systems to solve complex reasoning tasks, these agent executions often fail to meet human standards, leading to errors that compromise the system's overall…

人工智能 · 计算机科学 2025-03-18 Yoo Yeon Sung , Hannah Kim , Dan Zhang