English
Related papers

Related papers: How to use and interpret activation patching

200 papers

Activation oracles aim to make the activations of other models legible to humans and yield promising results compared to white-box interpretability techniques. However, uncertainty quantification (UQ) for the natural-language outputs of…

Computation and Language · Computer Science 2026-05-26 Federico Torrielli , Peter Schneider-Kamp , Lukas Galke Poech

Various techniques are used to detect the presence of charged particles stored in electromagnetic traps, their energy, their mass, or their internal states. Detection methods can rely on the variation of the number of trapped particles…

Atomic Physics · Physics 2017-08-23 Martina Knoop

New technologies have led to vast troves of large and complex datasets across many scientific domains and industries. People routinely use machine learning techniques to not only process, visualize, and make predictions from this big data,…

Machine Learning · Statistics 2023-08-04 Genevera I. Allen , Luqin Gan , Lili Zheng

This work presents an adaptive activation method for neural networks that exploits the interdependency of features. Each pixel, node, and layer is assigned with a polynomial activation function, whose coefficients are provided by an…

Computer Vision and Pattern Recognition · Computer Science 2018-11-22 Jinhyeok Jang , Jaehong Kim , Jaeyeon Lee , Seungjoon Yang

Inspired by biological neurons, the activation functions play an essential part in the learning process of any artificial neural network commonly used in many real-world problems. Various activation functions have been proposed in the…

Machine Learning · Computer Science 2022-12-29 Ameya D. Jagtap , George Em Karniadakis

We investigate the capacity of a flat partially reactive patch of arbitrary shape to trap independent particles that undergo steady-state diffusion in the three-dimensional space. We focus on the total flux of particles onto the patch that…

Chemical Physics · Physics 2026-02-02 Denis S. Grebenkov , Raphael Maurette

The debate around the interpretability of attention mechanisms is centered on whether attention scores can be used as a proxy for the relative amounts of signal carried by sub-components of data. We propose to study the interpretability of…

Machine Learning · Computer Science 2022-07-27 Jonathan Haab , Nicolas Deutschmann , Maria Rodríguez Martínez

Mechanistic interpretability aims to understand neural networks by identifying which learned features mediate specific behaviors. Attribution graphs reveal these feature pathways, but interpreting them requires extensive manual analysis --…

Computation and Language · Computer Science 2025-11-11 Giuseppe Birardi

This work presents a detailed covariance and correlation matrix analysis for experimentally measured cross sections obtained using the activation technique. Both statistical and systematic contributions to the covariance matrix were…

Nuclear Theory · Physics 2026-04-01 Tanmoy Bar

We propose an optomechanical setup where the activation of entanglement through the pre-availability of non-classical correlations can be demonstrated experimentally. We analyse the conditions under which the scheme is successful and relate…

Quantum Physics · Physics 2011-12-21 Laura Mazzola , Mauro Paternostro

SHAP is a popular method for measuring variable importance in machine learning models. In this paper, we study the algorithm used to estimate SHAP scores and outline its connection to the functional ANOVA decomposition. We use this…

Methodology · Statistics 2022-11-14 Andrew Herren , P. Richard Hahn

For machine learning models to be most useful in numerous sociotechnical systems, many have argued that they must be human-interpretable. However, despite increasing interest in interpretability, there remains no firm consensus on how to…

Machine Learning · Computer Science 2021-02-03 Andrew Slavin Ross , Nina Chen , Elisa Zhao Hang , Elena L. Glassman , Finale Doshi-Velez

Explainability is motivated by the lack of transparency of black-box Machine Learning approaches, which do not foster trust and acceptance of Machine Learning algorithms. This also happens in the Predictive Process Monitoring field, where…

Artificial Intelligence · Computer Science 2025-07-25 Williams Rizzi , Marco Comuzzi , Chiara Di Francescomarino , Chiara Ghidini , Suhwan Lee , Fabrizio Maria Maggi , Alexander Nolte

In the study of neural network interpretability, there is growing evidence to suggest that relevant features are encoded across many neurons in a distributed fashion. Making sense of these distributed representations without knowledge of…

Machine Learning · Computer Science 2025-01-28 Kyle Reing , Greg Ver Steeg , Aram Galstyan

The fields of explainable AI and mechanistic interpretability aim to uncover the internal structure of neural networks, with circuit discovery as a central tool for understanding model computations. Existing approaches, however, rely on…

Machine Learning · Computer Science 2026-03-05 Elena Golimblevskaia , Aakriti Jain , Bruno Puri , Ammar Ibrahim , Wojciech Samek , Sebastian Lapuschkin

Actionable Knowledge Discovery (AKD) is a crucial aspect of data mining that is gaining popularity and being applied in a wide range of domains. This is because AKD can extract valuable insights and information, also known as knowledge,…

Machine Learning · Computer Science 2023-01-24 Sayed Erfan Arefin

One of the main challenges in mechanistic interpretability is circuit discovery, determining which parts of a model perform a given task. We build on the Mechanistic Interpretability Benchmark (MIB) and propose three key improvements to…

Computation and Language · Computer Science 2025-10-31 Yaniv Nikankin , Dana Arad , Itay Itzhak , Anja Reusch , Adi Simhi , Gal Kesten-Pomeranz , Yonatan Belinkov

With machine learning models being increasingly used to aid decision making even in high-stakes domains, there has been a growing interest in developing interpretable models. Although many supposedly interpretable models have been proposed,…

Artificial Intelligence · Computer Science 2021-08-17 Forough Poursabzi-Sangdeh , Daniel G. Goldstein , Jake M. Hofman , Jennifer Wortman Vaughan , Hanna Wallach

This note explores probabilistic sampling weighted by uncertainty in active learning. This method has been previously used and authors have tangentially remarked on its efficacy. The scheme has several benefits: (1) it is computationally…

Machine Learning · Computer Science 2019-09-12 Vinay Jethava

We provide a novel notion of what it means to be interpretable, looking past the usual association with human understanding. Our key insight is that interpretability is not an absolute concept and so we define it relative to a target model,…

Artificial Intelligence · Computer Science 2018-10-30 Amit Dhurandhar , Vijay Iyengar , Ronny Luss , Karthikeyan Shanmugam