English
Related papers

Related papers: Minimum Levels of Interpretability for Artificial …

200 papers

Artificial Intelligence (AI) has continued to achieve tremendous success in recent times. However, the decision logic of these frameworks is often not transparent, making it difficult for stakeholders to understand, interpret or explain…

Machine Learning · Computer Science 2025-01-20 Fuseini Mumuni , Alhassan Mumuni

Mechanistic interpretability (MI) is an emerging framework for interpreting neural networks. Given a task and model, MI aims to discover a succinct algorithmic process, an interpretation, that explains the model's decision process on that…

Machine Learning · Computer Science 2026-04-01 Alan Sun , Mariya Toneva

How to attribute responsibility for autonomous artificial intelligence (AI) systems' actions has been widely debated across the humanities and social science disciplines. This work presents two experiments ($N$=200 each) that measure…

Computers and Society · Computer Science 2021-02-02 Gabriel Lima , Nina Grgić-Hlača , Meeyoung Cha

Automated interpretability systems aim to reduce the need for human labor and scale analysis to increasingly large models and diverse tasks. Recent efforts toward this goal leverage large language models (LLMs) at increasing levels of…

Artificial Intelligence · Computer Science 2026-03-23 Tal Haklay , Nikhil Prakash , Sana Pandey , Antonio Torralba , Aaron Mueller , Jacob Andreas , Tamar Rott Shaham , Yonatan Belinkov

The problem of human trust in artificial intelligence is one of the most fundamental problems in applied machine learning. Our processes for evaluating AI trustworthiness have substantial ramifications for ML's impact on science, health,…

Machine Learning · Computer Science 2022-02-14 Max W. Shen

Machine learning is an important tool for decision making, but its ethical and responsible application requires rigorous vetting of its interpretability and utility: an understudied problem, particularly for natural language processing…

Artificial Intelligence · Computer Science 2019-06-11 Shi Feng , Jordan Boyd-Graber

The increasing demand in artificial intelligence (AI) for models that are both effective and explainable is critical in domains where safety and trust are paramount. In this study, we introduce MIRA, a transparent and interpretable…

Artificial Intelligence · Computer Science 2024-10-18 Sarah Seifi , Tobias Sukianto , Cecilia Carbonelli , Lorenzo Servadei , Robert Wille

Responsible AI demands systems whose behavioral tendencies can be effectively measured, audited, and adjusted to prevent inadvertently nudging users toward risky decisions or embedding hidden biases in risk aversion. As language models…

Artificial Intelligence · Computer Science 2025-10-10 Ali Mazyaki , Mohammad Naghizadeh , Samaneh Ranjkhah Zonouzaghi , Amirhossein Farshi Sotoudeh

Existing approaches for generating human-aware agent behaviors have considered different measures of interpretability in isolation. Further, these measures have been studied under differing assumptions, thus precluding the possibility of…

Artificial Intelligence · Computer Science 2021-04-23 Sarath Sreedharan , Anagha Kulkarni , David E. Smith , Subbarao Kambhampati

As the complexity of AI systems and their interactions with the world increases, generating explanations for their behaviour is important for safely deploying AI. For agents, the most natural abstractions for predicting behaviour attribute…

Artificial Intelligence · Computer Science 2025-06-05 Alexis Bellot , Jonathan Richens , Tom Everitt

In this paper we describe moral quasi-dilemmas (MQDs): situations similar to moral dilemmas, but in which an agent is unsure whether exploring the plan space or the world may reveal a course of action that satisfies all moral requirements.…

Artificial Intelligence · Computer Science 2018-07-10 Daniel Kasenberg , Vasanth Sarathy , Thomas Arnold , Matthias Scheutz , Tom Williams

As AI agents built on large language models (LLMs) become increasingly embedded in society, issues of coordination, control, delegation, and accountability are entangled with concerns over their reliability. To design and implement LLM…

Computers and Society · Computer Science 2025-12-09 R. Patrick Xian , Garry A. Gabison , Ahmed Alaa , Christoph Riedl , Grigorios G. Chrysos

Explainability in AI and ML models is critical for fostering trust, ensuring accountability, and enabling informed decision making in high stakes domains. Yet this objective is often unmet in practice. This paper proposes a general purpose…

Statistical Finance · Quantitative Finance 2025-09-03 N. Jean , G. Le Pera

Artificial Intelligence (AI) has a tremendous impact on the unexpected growth of technology in almost every aspect. AI-powered systems are monitoring and deciding about sensitive economic and societal issues. The future is towards…

Machine Learning · Computer Science 2022-06-14 Ioannis Mollas , Nick Bassiliades , Grigorios Tsoumakas

This article proposes a new integration of linguistic anthropology and machine learning (ML) around convergent interests in both the underpinnings of language and making language technologies more socially responsible. While linguistic…

Computers and Society · Computer Science 2024-11-11 Graham M. Jones , Shai Satran , Arvind Satyanarayan

Is it possible to evaluate the moral cognition of complex artificial agents? In this work, we take a look at one aspect of morality: `doing the right thing for the right reasons.' We propose a behavior-based analysis of artificial moral…

Artificial Intelligence · Computer Science 2023-05-30 Yiran Mao , Madeline G. Reinecke , Markus Kunesch , Edgar A. Duéñez-Guzmán , Ramona Comanescu , Julia Haas , Joel Z. Leibo

As artificial intelligence systems increasingly inform high-stakes decisions across sectors, transparency has become foundational to responsible and trustworthy AI implementation. Leveraging our role as a leading institute in advancing AI…

Machine Learning · Computer Science 2025-08-01 Dhanesh Ramachandram , Himanshu Joshi , Judy Zhu , Dhari Gandhi , Lucas Hartman , Ananya Raval

We consider two fundamental and related issues currently faced by Artificial Intelligence (AI) development: the lack of ethics and interpretability of AI decisions. Can interpretable AI decisions help to address ethics in AI? Using a…

Artificial Intelligence · Computer Science 2021-09-21 Jean-Marie John-Mathews

The concepts of blameworthiness and wrongness are of fundamental importance in human moral life. But to what extent are humans disposed to blame artificially intelligent agents, and to what extent will they judge their actions to be morally…

Computers and Society · Computer Science 2021-02-09 Michael T. Stuart , Markus Kneer

AI practitioners increasingly use large language model (LLM) agents in compound AI systems to solve complex reasoning tasks, these agent executions often fail to meet human standards, leading to errors that compromise the system's overall…

Artificial Intelligence · Computer Science 2025-03-18 Yoo Yeon Sung , Hannah Kim , Dan Zhang