English

Understanding Task Representations in Neural Networks via Bayesian Ablation

Machine Learning 2026-04-07 v2 Artificial Intelligence

Abstract

Neural networks are powerful tools for cognitive modeling due to their flexibility and emergent properties. However, interpreting their learned representations remains challenging due to their sub-symbolic semantics. In this work, we introduce a novel probabilistic framework for interpreting latent task representations in neural networks. Inspired by Bayesian inference, our approach defines a distribution over representational units to infer their causal contributions to task performance. Using ideas from information theory, we propose a suite of tools and metrics to illuminate key model properties, including representational distributedness, manifold complexity, and polysemanticity.

Keywords

Cite

@article{arxiv.2505.13742,
  title  = {Understanding Task Representations in Neural Networks via Bayesian Ablation},
  author = {Andrew Nam and Declan Campbell and Thomas Griffiths and Jonathan Cohen and Sarah-Jane Leslie},
  journal= {arXiv preprint arXiv:2505.13742},
  year   = {2026}
}

Comments

Accepted at CLeaR 2026 (5th Conference on Causal Learning and Reasoning). 13 pages, 3 figures, plus appendix

R2 v1 2026-07-01T02:23:30.744Z