English
Related papers

Related papers: How to use and interpret activation patching

200 papers

In many real-world applications, the reward function is too complex to be manually specified. In such cases, reward functions must instead be learned from human feedback. Since the learned reward may fail to represent user preferences, it…

Machine Learning · Computer Science 2022-03-28 Erik Jenner , Adam Gleave

Transformers are widely used in natural language processing, where they consistently achieve state-of-the-art performance. This is mainly due to their attention-based architecture, which allows them to model rich linguistic relations…

Computation and Language · Computer Science 2022-11-29 Nikolaos Mylonas , Ioannis Mollas , Grigorios Tsoumakas

Humor is a magnetic component in everyday human interactions and communications. Computationally modeling humor enables NLP systems to entertain and engage with users. We investigate the effectiveness of prompting, a new transfer learning…

Computation and Language · Computer Science 2022-10-26 Junze Li , Mengjie Zhao , Yubo Xie , Antonis Maronikolakis , Pearl Pu , Hinrich Schütze

The lack of interpretability is an inevitable problem when using neural network models in real applications. In this paper, an explainable neural network based on generalized additive models with structured interactions (GAMI-Net) is…

Machine Learning · Statistics 2021-06-03 Zebin Yang , Aijun Zhang , Agus Sudjianto

Because of the pervasive usage of Neural Networks in human sensitive applications, their interpretability is becoming an increasingly important topic in machine learning. In this work we introduce a simple way to interpret the output…

Machine Learning · Computer Science 2021-02-08 Stefano Zamuner , Paolo De Los Rios

Decisions by Machine Learning (ML) models have become ubiquitous. Trusting these decisions requires understanding how algorithms take them. Hence interpretability methods for ML are an active focus of research. A central problem in this…

Machine Learning · Computer Science 2019-01-25 Philipp Schmidt , Felix Biessmann

We often desire our models to be interpretable as well as accurate. Prior work on optimizing models for interpretability has relied on easy-to-quantify proxies for interpretability, such as sparsity or the number of operations required. In…

Machine Learning · Statistics 2018-11-01 Isaac Lage , Andrew Slavin Ross , Been Kim , Samuel J. Gershman , Finale Doshi-Velez

We review some applications of entanglement to improve quantum measurements and communication, with the main focus on the optical implementation of quantum information processing. The evolution of continuos variable entangled states in…

Quantum Physics · Physics 2007-05-23 M. G. A. Paris , G. M. D'Ariano , P. Lo Presti , P. Perinotti

We explain the meaning of dynamical manipulation, and we illustrate its mechanism by using a system composed of a charged particle in a Penning trap. It is shown that by means of appropriate electric shocks (delta-like pulses) applied to…

Quantum Physics · Physics 2007-05-23 David J. Fernandez C

Despite recent progress in artificial intelligence and machine learning, many state-of-the-art methods suffer from a lack of explainability and transparency. The ability to interpret the predictions made by machine learning models and…

Machine Learning · Computer Science 2021-11-10 Zihan Wang , Jialin Lu , Oliver Snow , Martin Ester

Interpretability methods aim to help users build trust in and understand the capabilities of machine learning models. However, existing approaches often rely on abstract, complex visualizations that poorly map to the task at hand or require…

Human-Computer Interaction · Computer Science 2021-07-12 Harini Suresh , Kathleen M. Lewis , John V. Guttag , Arvind Satyanarayan

After a brief introduction into the general importance of polarization observables for the analysis of a reaction, the basic density matrix formalism for the description of polarization phenomena is outlined and illustrated by explicit…

Nuclear Theory · Physics 2015-05-13 Hartmuth Arenhoevel

This paper presents an automated approach for interpretable feature recommendation for solving signal data analytics problems. The method has been tested by performing experiments on datasets in the domain of prognostics where…

Machine Learning · Statistics 2017-11-07 Snehasis Banerjee , Tanushyam Chattopadhyay , Ayan Mukherjee

Attitudes about artificial intelligence and machine learning are recent victims of endemic misunderstanding; given our increasing reliance on these technologies, the need for widespread understanding and confidence in their use is…

Graphics · Computer Science 2026-05-04 Bokang Wang , Yingxuan Liao , Leah Lee , Jack Wesson , Anlan Yang , Ruizi Wang , Yigang Wen

Mechanistic interpretability aims to understand the internal mechanisms learned by neural networks. Despite recent progress toward this goal, it remains unclear how best to decompose neural network parameters into mechanistic components. We…

Machine Learning · Computer Science 2025-02-11 Dan Braun , Lucius Bushnaq , Stefan Heimersheim , Jake Mendel , Lee Sharkey

State of the art machine learning algorithms are highly optimized to provide the optimal prediction possible, naturally resulting in complex models. While these models often outperform simpler more interpretable models by order of…

Machine Learning · Statistics 2016-11-24 Yotam Hechtlinger

Active learning identifies data points to label that are expected to be the most useful in improving a supervised model. Opportunistic active learning incorporates active learning into interactive tasks that constrain possible queries…

Computation and Language · Computer Science 2018-08-31 Aishwarya Padmakumar , Peter Stone , Raymond J. Mooney

This chapter revisits the concept of excitability, a basic system property of neurons. The focus is on excitable systems regarded as behaviors rather than dynamical systems. By this we mean open systems modulated by specific interconnection…

Neurons and Cognition · Quantitative Biology 2017-04-18 Rodolphe Sepulchre , Guillaume Drion , Alessio Franci

A foremost challenge in modern network science is the inverse problem of reconstruction (inference) of coupling equations and network topology from the measurements of the network dynamics. Of particular interest are the methods that can…

Chaotic Dynamics · Physics 2019-10-31 Isao T. Tokuda , Zoran Levnajic , Kazuyoshi Ishimura

How interpretable are the features of leading vision models? The question is increasingly pressing as these models move from research benchmarks into high-stakes deployments, yet existing methods cannot answer it reliably. We close this gap…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Julien Colin , Lore Goetschalckx , Nuria Oliver , Thomas Serre