中文
相关论文

相关论文: On the (In)fidelity and Sensitivity for Explanatio…

200 篇论文

We present an interpretable companion model for any pre-trained black-box classifiers. The idea is that for any input, a user can decide to either receive a prediction from the black-box model, with high accuracy but no explanations, or…

机器学习 · 统计学 2020-02-12 Danqing Pan , Tong Wang , Satoshi Hara

Estimating the fidelity with a target state is important in quantum information tasks. Many fidelity estimation techniques present a suitable measurement scheme to perform the estimation. In contrast, we present techniques that allow the…

量子物理 · 物理学 2024-07-12 Akshay Seshadri , Martin Ringbauer , Jacob Spainhour , Thomas Monz , Stephen Becker

We build on abduction-based explanations for ma-chine learning and develop a method for computing local explanations for neural network models in natural language processing (NLP). Our explanations comprise a subset of the words of the…

Large Language Models (LLMs) offer natural language explanations as an alternative to feature attribution methods for model interpretability. However, despite their plausibility, they may not reflect the model's true reasoning faithfully.…

计算与语言 · 计算机科学 2025-12-29 Kerem Zaman , Shashank Srivastava

The sensitivities revealed by a sensitivity analysis of a probabilistic network typically depend on the entered evidence. For a real-life network therefore, the analysis is performed a number of times, with different evidence. Although…

人工智能 · 计算机科学 2012-07-19 Silja Renooij , Linda C. van der Gaag

We present a new procedure for conducting a sensitivity analysis in matched observational studies. For any candidate test statistic, the approach defines tilted modifications dependent upon the proposed strength of unmeasured confounding.…

统计方法学 · 统计学 2025-03-14 Colin B. Fogarty

Probabilistic confidence metrics are increasingly adopted as proxies for reasoning quality in Best-of-N selection, under the assumption that higher confidence reflects higher reasoning fidelity. In this work, we challenge this assumption by…

人工智能 · 计算机科学 2026-01-21 Hojin Kim , Jaehyung Kim

The objective of this paper is to assess the quality of explanation heatmaps for image classification tasks. To assess the quality of explainability methods, we approach the task through the lens of accuracy and stability. In this work, we…

计算机视觉与模式识别 · 计算机科学 2025-09-15 Lassi Raatikainen , Esa Rahtu

We study a worst-case approach to measure the sensitivity to model misspecification in the performance analysis of stochastic systems. The situation of interest is when only minimal parametric information is available on the form of the…

概率论 · 数学 2015-07-14 Henry Lam

In recent years, an abundance of feature attribution methods for explaining neural networks have been developed. Especially in the field of computer vision, many methods for generating saliency maps providing pixel attributions exist.…

计算机视觉与模式识别 · 计算机科学 2022-07-06 Yannik Mahlau , Christian Nolde

Fairness metrics are a core tool in the fair machine learning literature (FairML), used to determine that ML models are, in some sense, "fair". Real-world data, however, are typically plagued by various measurement biases and other violated…

机器学习 · 计算机科学 2024-10-16 Jake Fawkes , Nic Fishman , Mel Andrews , Zachary C. Lipton

With the availability of large databases and recent improvements in deep learning methodology, the performance of AI systems is reaching or even exceeding the human level on an increasing number of complex tasks. Impressive examples of this…

人工智能 · 计算机科学 2017-08-29 Wojciech Samek , Thomas Wiegand , Klaus-Robert Müller

Understanding why machine learning models behave the way they do empowers both system designers and end-users in many ways: in model selection, feature engineering, in order to trust and act upon the predictions, and in more intuitive user…

机器学习 · 统计学 2016-06-20 Marco Tulio Ribeiro , Sameer Singh , Carlos Guestrin

Interpretability is the study of explaining models in understandable terms to humans. At present, interpretability is divided into two paradigms: the intrinsic paradigm, which believes that only models designed to be explained can be…

机器学习 · 计算机科学 2024-11-14 Andreas Madsen , Himabindu Lakkaraju , Siva Reddy , Sarath Chandar

We study the robustness of Bayesian persuasion to uncertainty about the receiver's preferences. We analyze two conceptually distinct notions: continuity, in which only the modeler lacks precise knowledge, but where the model's predictions…

理论经济学 · 经济学 2026-05-28 Ronen Gradwohl , Fengming Hu , Rann Smorodinsky

A frequentist definition of sensitivity of a search for new phenomena is discussed, that has several useful properties. It is based on completely standard concepts, is generally applicable, and has a very clear interpretation. It is…

数据分析、统计与概率 · 物理学 2007-05-23 Giovanni Punzi

A key objective of decomposition analysis is to identify a factor (the 'mediator') contributing to disparities in an outcome between social groups. In decomposition analysis, a scholarly interest often centers on estimating how much the…

统计方法学 · 统计学 2022-05-27 Soojin Park , Suyeon Kang , Chioun Lee , Shujie Ma

Faithful explanations are essential for machine learning models in high-stakes applications. Inherently interpretable models are well-suited for these applications because they naturally provide faithful explanations by revealing their…

机器学习 · 计算机科学 2025-02-28 Chudi Zhong , Panyu Chen , Cynthia Rudin

As neural networks become more popular, the need for accompanying uncertainty estimates increases. There are currently two main approaches to test the quality of these estimates. Most methods output a density. They can be compared by…

机器学习 · 统计学 2024-06-05 Laurens Sluijterman , Eric Cator , Tom Heskes

Model steering, which involves intervening on hidden representations at inference time, has emerged as a lightweight alternative to finetuning for precisely controlling large language models. While steering efficacy has been widely studied,…

机器学习 · 计算机科学 2026-02-09 Navita Goyal , Hal Daumé