English
Related papers

Related papers: How to use and interpret activation patching

200 papers

In this survey we present different approaches that allow an intelligent agent to explore autonomous its environment to gather information and learn multiple tasks. Different communities proposed different solutions, that are in many cases,…

Artificial Intelligence · Computer Science 2014-03-07 Manuel Lopes , Luis Montesano

Activation steering methods in large language models (LLMs) have emerged as an effective way to perform targeted updates to enhance generated language without requiring large amounts of adaptation data. We ask whether the features…

Computation and Language · Computer Science 2025-11-05 Masha Fedzechkina , Eleonora Gualdoni , Sinead Williamson , Katherine Metcalf , Skyler Seto , Barry-John Theobald

We present a method for diagnosing interpretation in neural networks by identifying an input subspace where a proposed interpretation is highly faithful. Our method is particularly useful for causal-abstraction-style interpretability, where…

Artificial Intelligence · Computer Science 2026-05-05 Li Puyin , Jiyuan Tan , Ahmad Jabbar , Thomas Icard , Atticus Geiger

The circuits framework in mechanistic interpretability aims to identify causally important sparse subgraphs of model components, typically evaluated by measuring necessity and sufficiency. We measure circuit reuse, the proportion of…

Computation and Language · Computer Science 2026-05-12 Michael Li , Nishant Subramani

Interpreting the internal activations of neural networks can produce more faithful explanations of their behavior, but is difficult due to the complex structure of activation space. Existing approaches to scalable interpretability use…

Artificial Intelligence · Computer Science 2025-12-18 Vincent Huang , Dami Choi , Daniel D. Johnson , Sarah Schwettmann , Jacob Steinhardt

Rollating walkers are popular mobility aids used by older adults to improve balance control. There is a need to automatically recognize the activities performed by walker users to better understand activity patterns, mobility issues and the…

Artificial Intelligence · Computer Science 2012-03-19 Farheen Omar , Mathieu Sinn , Jakub Truszkowski , Pascal Poupart , James Tung , Allen Caine

Understanding AI systems' inner workings is critical for ensuring value alignment and safety. This review explores mechanistic interpretability: reverse engineering the computational mechanisms and representations learned by neural networks…

Artificial Intelligence · Computer Science 2024-08-27 Leonard Bereska , Efstratios Gavves

We propose a unified, human-readable, machine-processable novel syntax/notation designed to comprehensively describe reactions, molecules and excitation states. Our notation resolves inconsistencies in existing data representations and…

Computational Physics · Physics 2025-04-15 Dan Andrei Ciubotaru , Michele Renda , Călin Alexa

In this paper, the problem of choosing the best allocation of excitations and measurements for the identification of a dynamic network is formally stated and analyzed. The best choice will be one that achieves the most accurate…

Optimization and Control · Mathematics 2021-09-21 Eduardo Mapurunga , Alexandre Sanfelici Bazanella

A common approach to interpreting spiking activity is based on identifying the firing fields---regions in physical or configuration spaces that elicit responses of neurons. Common examples include hippocampal place cells that fire at…

Neurons and Cognition · Quantitative Biology 2021-08-10 D. Akhtiamov , A. G. Cohn , Y. Dabaghian

Along with the great success of deep neural networks, there is also growing concern about their black-box nature. The interpretability issue affects people's trust on deep learning systems. It is also related to many ethical problems, e.g.,…

Machine Learning · Computer Science 2022-02-01 Yu Zhang , Peter Tiňo , Aleš Leonardis , Ke Tang

We present a new multi-objective optimization approach for synthesizing interpretations that "explain" the behavior of black-box machine learning models. Constructing human-understandable interpretations for black-box models often requires…

Machine Learning · Computer Science 2021-08-18 Hazem Torfah , Shetal Shah , Supratik Chakraborty , S. Akshay , Sanjit A. Seshia

Entanglement is a fundamental property of quantum systems, essential for non-trivial quantum programs. Identifying when qubits become entangled is critical for circuit optimization, and for arguing for the correctness of quantum algorithms.…

Quantum Physics · Physics 2025-08-15 Aske Nord Raahauge , Martin Bom Marchioro , Rasmus Ross Nylandsted

Mechanistic interpretability is an emerging diagnostic approach for neural models that has gained traction in broader natural language processing domains. This paradigm aims to provide attribution to components of neural systems where…

Information Retrieval · Computer Science 2025-01-20 Andrew Parry , Catherine Chen , Carsten Eickhoff , Sean MacAvaney

In this paper, we extend to polarization the method we have recently employed to treat spin. We are led to a generalization of its treatment. Thus, we are able to connect its matrix treatment to first principles, and we obtain the most…

Quantum Physics · Physics 2007-05-23 Habatwa V. Mweene

In-circuit impedance provides key information for many EMC applications. The inductive coupling approach is a promising method for in-circuit impedance measurement because its measurement setups have no direct electrical contact with the…

Instrumentation and Detectors · Physics 2022-04-05 Zhenyu Zhao , Fei Fan , Huamin Jie , Quqin Sun , Pengfei Tu , Wensong Wang , Kye Yak See

Attention mechanisms have recently boosted performance on a range of NLP tasks. Because attention layers explicitly weight input components' representations, it is also often assumed that attention can be used to identify information that…

Computation and Language · Computer Science 2019-06-11 Sofia Serrano , Noah A. Smith

Post-hoc model-agnostic interpretation methods such as partial dependence plots can be employed to interpret complex machine learning models. While these interpretation methods can be applied regardless of model complexity, they can produce…

Machine Learning · Statistics 2022-01-24 Christoph Molnar , Giuseppe Casalicchio , Bernd Bischl

The data mining process consists of a series of steps ranging from data cleaning, data selection and transformation, to pattern evaluation and visualization. One of the central problems in data mining is to make the mined patterns or…

Databases · Computer Science 2007-05-23 Zengyou He , Xiaofei Xu , Shengchun Deng

Machine learning transparency calls for interpretable explanations of how inputs relate to predictions. Feature attribution is a way to analyze the impact of features on predictions. Feature interactions are the contextual dependence…

Machine Learning · Statistics 2020-06-22 Michael Tsang , Sirisha Rambhatla , Yan Liu