English
Related papers

Related papers: How to use and interpret activation patching

200 papers

Neural networks are growing more capable on their own, but we do not understand their neural mechanisms. Understanding these mechanisms' decision-making processes, or mechanistic interpretability, enables (1) accountability and control in…

Computation and Language · Computer Science 2026-03-02 Mason Kadem , Rong Zheng

Explainable AI (XAI) methods reveal which features influence model predictions, yet provide limited means for practitioners to act on these explanations. Activation steering of components identified via XAI offers a path toward actionable…

Artificial Intelligence · Computer Science 2026-05-27 Tobias Labarta , Maximilian Dreyer , Katharina Weitz , Wojciech Samek , Sebastian Lapuschkin

To study implementations and optimisations of interaction net systems we propose a calculus to allow us to reason about nets, a concrete data-structure that is in close correspondence with the calculus, and a low-level language to create…

Logic in Computer Science · Computer Science 2015-05-28 Abubakar Hassan , Ian Mackie , Shinya Sato

As machine learning systems become ubiquitous, there has been a surge of interest in interpretable machine learning: systems that provide explanation for their outputs. These explanations are often used to qualitatively assess other…

Machine Learning · Statistics 2017-03-06 Finale Doshi-Velez , Been Kim

With the broader and highly successful usage of machine learning in industry and the sciences, there has been a growing demand for Explainable AI. Interpretability and explanation methods for gaining a better understanding about the problem…

Machine Learning · Computer Science 2021-02-26 Wojciech Samek , Grégoire Montavon , Sebastian Lapuschkin , Christopher J. Anders , Klaus-Robert Müller

Mechanistic interpretability is the program of explaining what AI systems are doing in terms of their internal mechanisms. I analyze some aspects of the program, along with setting out some concrete challenges and assessing progress to…

Artificial Intelligence · Computer Science 2025-01-28 David J. Chalmers

Neural network interpretability is a vital component for applications across a wide variety of domains. In such cases it is often useful to analyze a network which has already been trained for its specific purpose. In this work, we develop…

Machine Learning · Computer Science 2019-11-19 Lawrence Phillips , Garrett Goh , Nathan Hodas

This paper argues that interpretability research in Artificial Intelligence (AI) is fundamentally ill-posed as existing definitions of interpretability fail to describe how interpretability can be formally tested or designed for. We posit…

Artificial Intelligence · Computer Science 2026-01-30 Pietro Barbiero , Mateo Espinosa Zarlenga , Francesco Giannini , Alberto Termine , Filippo Bonchi , Mateja Jamnik , Giuseppe Marra

Recent approaches in robotics follow the insight that perception is facilitated by interaction with the environment. These approaches are subsumed under the term of Interactive Perception (IP). It provides the following benefits: (i)…

Probing (or diagnostic classification) has become a popular strategy for investigating whether a given set of intermediate features is present in the representations of neural models. Probing studies may have misleading results, but various…

Machine Learning · Computer Science 2021-10-01 Deborah Ferreira , Julia Rozanova , Mokanarangan Thayaparan , Marco Valentino , André Freitas

In recent years the term \textit{active wetting} has gained some traction in works describing, analyzing and modeling a wide variety of wetting phenomena, for instance, in the contexts of biomolecular condensates, of cell layers or cell…

Soft Condensed Matter · Physics 2026-02-12 Uwe Thiele

Soft prompts have been popularized as a cheap and easy way to improve task-specific LLM performance beyond few-shot prompts. Despite their origin as an automated prompting method, however, soft prompts and other trainable prompts remain a…

Machine Learning · Computer Science 2025-04-04 Oam Patel , Jason Wang , Nikhil Shivakumar Nayak , Suraj Srinivas , Himabindu Lakkaraju

Two computational models to be used as tools for experimental research on the retinal implant are presented. In the first model, the electric field produced by a multi-electrode array in a uniform retina is calculated. In the second model,…

Neurons and Cognition · Quantitative Biology 2010-12-30 Erich W. Schmid , Robert Wilke

As AI systems are used in high-stakes applications, ensuring interpretability is crucial. Mechanistic Interpretability (MI) aims to reverse-engineer neural networks by extracting human-understandable algorithms to explain their behavior.…

Machine Learning · Computer Science 2025-03-03 Maxime Méloux , Silviu Maniu , François Portet , Maxime Peyrard

There is a need of ensuring machine learning models that are interpretable. Higher interpretability of the model means easier comprehension and explanation of future predictions for end-users. Further, interpretable machine learning models…

Machine Learning · Computer Science 2020-08-17 Gregor Stiglic , Primoz Kocbek , Nino Fijacko , Marinka Zitnik , Katrien Verbert , Leona Cilar

In neural networks literature, there is a strong interest in identifying and defining activation functions which can improve neural network performance. In recent years there has been a renovated interest of the scientific community in…

Machine Learning · Computer Science 2021-03-01 Andrea Apicella , Francesco Donnarumma , Francesco Isgrò , Roberto Prevete

Mechanistic interpretability (MI) is an emerging framework for interpreting neural networks. Given a task and model, MI aims to discover a succinct algorithmic process, an interpretation, that explains the model's decision process on that…

Machine Learning · Computer Science 2026-04-01 Alan Sun , Mariya Toneva

When explaining the decisions of deep neural networks, simple stories are tempting but dangerous. Especially in computer vision, the most popular explanation approaches give a false sense of comprehension to its users and provide an overly…

Machine Learning · Computer Science 2021-09-17 Matthias Kirchler , Martin Graf , Marius Kloft , Christoph Lippert

Several different types of statistical interaction are defined and distinguished, primarily on the basis of the nature of the factors defining the interaction. Illustrative examples, mostly epidemiological, are given. The emphasis is…

Applications · Statistics 2009-09-29 Amy Berrington de González , D. R. Cox

Recent years have seen a boom in interest in machine learning systems that can provide a human-understandable rationale for their predictions or decisions. However, exactly what kinds of explanation are truly human-interpretable remains…

Machine Learning · Computer Science 2019-08-30 Isaac Lage , Emily Chen , Jeffrey He , Menaka Narayanan , Been Kim , Sam Gershman , Finale Doshi-Velez