English

What Happens to a Dataset Transformed by a Projection-based Concept Removal Method?

Computation and Language 2024-03-26 v1 Artificial Intelligence

Abstract

We investigate the behavior of methods that use linear projections to remove information about a concept from a language representation, and we consider the question of what happens to a dataset transformed by such a method. A theoretical analysis and experiments on real-world and synthetic data show that these methods inject strong statistical dependencies into the transformed datasets. After applying such a method, the representation space is highly structured: in the transformed space, an instance tends to be located near instances of the opposite label. As a consequence, the original labeling can in some cases be reconstructed by applying an anti-clustering method.

Cite

@article{arxiv.2403.16142,
  title  = {What Happens to a Dataset Transformed by a Projection-based Concept Removal Method?},
  author = {Richard Johansson},
  journal= {arXiv preprint arXiv:2403.16142},
  year   = {2024}
}
R2 v1 2026-06-28T15:31:39.647Z