中文
相关论文

相关论文: Meaningfully Debugging Model Mistakes using Concep…

200 篇论文

Counterfactual statements, which describe events that did not or cannot take place, are beneficial to numerous NLP applications. Hence, we consider the problem of counterfactual detection (CFD) and seek to enhance the CFD models. Previous…

计算与语言 · 计算机科学 2024-10-01 Thong Nguyen , Truc-My Nguyen

Causal concept effect estimation is gaining increasing interest in the field of interpretable machine learning. This general approach explains the behaviors of machine learning models by estimating the causal effect of human-understandable…

机器学习 · 计算机科学 2024-11-15 Jifan Gao , Guanhua Chen

Learning rewards from human behaviour or feedback is a promising approach to aligning AI systems with human values but fails to consistently extract correct reward functions. Interpretability tools could enable users to understand and…

人工智能 · 计算机科学 2024-10-16 Jan Wehner , Frans Oliehoek , Luciano Cavalcante Siebert

There has been a growing interest in model-agnostic methods that can make deep learning models more transparent and explainable to a user. Some researchers recently argued that for a machine to achieve a certain degree of human-level…

人工智能 · 计算机科学 2021-06-09 Yu-Liang Chou , Catarina Moreira , Peter Bruza , Chun Ouyang , Joaquim Jorge

Verification and validation of cyber-physical systems (CPS) via large-scale simulation often surface failures that are hard to interpret, especially when triggered by interactions between continuous and discrete behaviors at specific events…

软件工程 · 计算机科学 2026-04-10 Zaid Ghazal , Hadiza Yusuf , Khouloud Gaaloul

Counterfactual explanations indicate the smallest change in input that can translate to a different outcome for a machine learning model. Counterfactuals have generated immense interest in high-stakes applications such as finance,…

机器学习 · 计算机科学 2025-03-12 Erfaun Noorani , Pasan Dissanayake , Faisal Hamman , Sanghamitra Dutta

Counterfactual Explanations are becoming a de-facto standard in post-hoc interpretable machine learning. For a given classifier and an instance classified in an undesired class, its counterfactual explanation corresponds to small…

机器学习 · 计算机科学 2024-01-17 Veronica Piccialli , Dolores Romero Morales , Cecilia Salvatore

Counterfactual instances are a powerful tool to obtain valuable insights into automated decision processes, describing the necessary minimal changes in the input space to alter the prediction towards a desired target. Most previous…

机器学习 · 计算机科学 2021-06-07 Robert-Florian Samoilescu , Arnaud Van Looveren , Janis Klaise

Machine learning (ML) methods have experienced significant growth in the past decade, yet their practical application in high-impact real-world domains has been hindered by their opacity. When ML methods are responsible for making critical…

机器学习 · 计算机科学 2025-07-11 Xiangyu Sun , Raquel Aoki , Kevin H. Wilson

Recent work on counterfactual visual explanations has contributed to making artificial intelligence models more explainable by providing visual perturbation to flip the prediction. However, these approaches neglect the causal relationships…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Yiran Qiao , Disheng Liu , Yiren Lu , Yu Yin , Mengnan Du , Jing Ma

Counterfactual explanations (CE) aim to reveal how small input changes flip a model's prediction, yet many methods modify more features than necessary, reducing clarity and actionability. We introduce \emph{COLA}, a model- and…

机器学习 · 计算机科学 2026-03-02 Lei You , Yijun Bian , Lele Cao

Both uncertainty estimation and interpretability are important factors for trustworthy machine learning systems. However, there is little work at the intersection of these two areas. We address this gap by proposing a novel method for…

Machine learning models that automate decision-making are increasingly used in consequential areas such as loan approvals, pretrial bail approval, and hiring. Unfortunately, most of these models are black boxes, i.e., they are unable to…

人工智能 · 计算机科学 2024-05-28 Sopam Dasgupta , Joaquín Arias , Elmer Salazar , Gopal Gupta

The increasing use of Machine Learning (ML) models to aid decision-making in high-stakes industries demands explainability to facilitate trust. Counterfactual Explanations (CEs) are ideally suited for this, as they can offer insights into…

机器学习 · 计算机科学 2025-02-20 Junqi Jiang , Luca Marzari , Aaryan Purohit , Francesco Leofante

Prediction failures of machine learning models often arise from deficiencies in training data, such as incorrect labels, outliers, and selection biases. However, such data points that are responsible for a given failure mode are generally…

机器学习 · 计算机科学 2022-11-11 Ryutaro Tanno , Melanie F. Pradier , Aditya Nori , Yingzhen Li

We present CounterfactualExplanations.jl: a package for generating Counterfactual Explanations (CE) and Algorithmic Recourse (AR) for black-box models in Julia. CE explain how inputs into a model need to change to yield specific model…

机器学习 · 计算机科学 2023-08-15 Patrick Altmeyer , Arie van Deursen , Cynthia C. S. Liem

Although machine learning (ML) models of AI achieve high performances in medicine, they are not free of errors. Empowering clinicians to identify incorrect model recommendations is crucial for engendering trust in medical AI. Explainable AI…

人工智能 · 计算机科学 2022-12-20 Isil Guzey , Ozlem Ucar , Nukhet Aladag Ciftdemir , Betul Acunas

In this work, we develop a technique to produce counterfactual visual explanations. Given a 'query' image $I$ for which a vision system predicts class $c$, a counterfactual visual explanation identifies how $I$ could change such that the…

机器学习 · 计算机科学 2019-06-12 Yash Goyal , Ziyan Wu , Jan Ernst , Dhruv Batra , Devi Parikh , Stefan Lee

To collaborate effectively with humans, language models must be able to explain their decisions in natural language. We study a specific type of self-explanation: self-generated counterfactual explanations (SCEs), where a model explains its…

机器学习 · 计算机科学 2025-09-12 Harry Mayne , Ryan Othniel Kearns , Yushi Yang , Andrew M. Bean , Eoin Delaney , Chris Russell , Adam Mahdi

Counterfactual explanations (CFEs) are a popular approach for interpreting machine learning predictions by identifying minimal feature changes that alter model outputs. However, in real-world settings, users often refine feasibility…

机器学习 · 计算机科学 2025-05-28 Christos Fragkathoulas , Evaggelia Pitoura