English

DePass: Unified Feature Attributing by Simple Decomposed Forward Pass

Computation and Language 2025-10-27 v2

Abstract

Attributing the behavior of Transformer models to internal computations is a central challenge in mechanistic interpretability. We introduce DePass, a unified framework for feature attribution based on a single decomposed forward pass. DePass decomposes hidden states into customized additive components, then propagates them with attention scores and MLP's activations fixed. It achieves faithful, fine-grained attribution without requiring auxiliary training. We validate DePass across token-level, model component-level, and subspace-level attribution tasks, demonstrating its effectiveness and fidelity. Our experiments highlight its potential to attribute information flow between arbitrary components of a Transformer model. We hope DePass serves as a foundational tool for broader applications in interpretability.

Cite

@article{arxiv.2510.18462,
  title  = {DePass: Unified Feature Attributing by Simple Decomposed Forward Pass},
  author = {Xiangyu Hong and Che Jiang and Kai Tian and Biqing Qi and Youbang Sun and Ning Ding and Bowen Zhou},
  journal= {arXiv preprint arXiv:2510.18462},
  year   = {2025}
}
R2 v1 2026-07-01T06:57:32.289Z