English

Mechanistic Interpretability of EEG Foundation Models via Sparse Autoencoders

Machine Learning 2026-05-25 v3 Human-Computer Interaction Neural and Evolutionary Computing

Abstract

EEG foundation models achieve state-of-the-art clinical performance, yet the internal computations driving their predictions remain opaque: a barrier to clinical trust. We apply TopK Sparse Autoencoders (SAEs) across three architecturally distinct EEG transformers: SleepFM, REVE, and LaBraM to extract sparse feature dictionaries from their embeddings. By grounding these features in a clinical taxonomy (abnormality, age, sex, and medication), we benchmark monosemanticity and entanglement across architectures. A single hyperparameter procedure, driven by an intrinsic dictionary health audit, transfers robustly across all three architectures. Via concept steering, we introduce a "target vs. off-target" probe area metric to quantify steering selectivity and reveal three operational regimes: selectively steerable, encoded but entangled, and non-encoded. This framework exposes critical representational failures: "wrecking-ball" interventions that collapse global model performance, and clinical entanglements, such as age-pathology confounding, where it is impossible to suppress one concept without corrupting the other. Finally, a spectral decoder maps these interventions back to the amplitude spectrum, translating latent manipulations into physiologically interpretable frequency signatures, such as pathological slow-wave suppression and α\alpha-band restoration.

Keywords

Cite

@article{arxiv.2605.13930,
  title  = {Mechanistic Interpretability of EEG Foundation Models via Sparse Autoencoders},
  author = {William Lehn-Schiøler and Magnus Ruud Kjær and Rahul Thapa and Magnus Guldberg Pedersen and Anton Mosquera Storgaard and Nick Williams and Radu Gatej and Tue Lehn-Schiøler and Andreas Brink-Kjær and Sadasivan Puthusserypady and Sándor Beniczky and James Zou and Lars Kai Hansen},
  journal= {arXiv preprint arXiv:2605.13930},
  year   = {2026}
}

Comments

Preprint. 14 pages, 7 figures, 4 tables

R2 v1 2026-07-22T07:10:53.208Z