English

LEACE: Perfect linear concept erasure in closed form

Machine Learning 2025-04-04 v4 Computation and Language Computers and Society

Abstract

Concept erasure aims to remove specified features from an embedding. It can improve fairness (e.g. preventing a classifier from using gender or race) and interpretability (e.g. removing a concept to observe changes in model behavior). We introduce LEAst-squares Concept Erasure (LEACE), a closed-form method which provably prevents all linear classifiers from detecting a concept while changing the embedding as little as possible, as measured by a broad class of norms. We apply LEACE to large language models with a novel procedure called "concept scrubbing," which erases target concept information from every layer in the network. We demonstrate our method on two tasks: measuring the reliance of language models on part-of-speech information, and reducing gender bias in BERT embeddings. Code is available at https://github.com/EleutherAI/concept-erasure.

Keywords

Cite

@article{arxiv.2306.03819,
  title  = {LEACE: Perfect linear concept erasure in closed form},
  author = {Nora Belrose and David Schneider-Joseph and Shauli Ravfogel and Ryan Cotterell and Edward Raff and Stella Biderman},
  journal= {arXiv preprint arXiv:2306.03819},
  year   = {2025}
}
R2 v1 2026-06-28T10:57:59.607Z