Training on Plausible Counterfactuals Removes Spurious Correlations
Machine Learning
2025-11-14 v8
Abstract
Plausible counterfactual explanations (p-CFEs) are perturbations that minimally modify inputs to change classifier decisions while remaining plausible under the data distribution. In this study, we demonstrate that classifiers can be trained on p-CFEs labeled with induced \emph{incorrect} target classes to classify unperturbed inputs with the original labels. While previous studies have shown that such learning is possible with adversarial perturbations, we extend this paradigm to p-CFEs. Interestingly, our experiments reveal that learning from p-CFEs is even more effective: the resulting classifiers achieve not only high in-distribution accuracy but also exhibit significantly reduced bias with respect to spurious correlations.
Cite
@article{arxiv.2505.16583,
title = {Training on Plausible Counterfactuals Removes Spurious Correlations},
author = {Shpresim Sadiku and Kartikeya Chitranshi and Hiroshi Kera and Sebastian Pokutta},
journal= {arXiv preprint arXiv:2505.16583},
year = {2025}
}