English

Generalizability vs. Counterfactual Explainability Trade-Off

Machine Learning 2025-05-30 v1

Abstract

In this work, we investigate the relationship between model generalization and counterfactual explainability in supervised learning. We introduce the notion of ε\varepsilon-valid counterfactual probability (ε\varepsilon-VCP) -- the probability of finding perturbations of a data point within its ε\varepsilon-neighborhood that result in a label change. We provide a theoretical analysis of ε\varepsilon-VCP in relation to the geometry of the model's decision boundary, showing that ε\varepsilon-VCP tends to increase with model overfitting. Our findings establish a rigorous connection between poor generalization and the ease of counterfactual generation, revealing an inherent trade-off between generalization and counterfactual explainability. Empirical results validate our theory, suggesting ε\varepsilon-VCP as a practical proxy for quantitatively characterizing overfitting.

Keywords

Cite

@article{arxiv.2505.23225,
  title  = {Generalizability vs. Counterfactual Explainability Trade-Off},
  author = {Fabiano Veglianti and Flavio Giorgi and Fabrizio Silvestri and Gabriele Tolomei},
  journal= {arXiv preprint arXiv:2505.23225},
  year   = {2025}
}

Comments

9 pages, 4 figures, plus appendix. arXiv admin note: text overlap with arXiv:2502.09193

R2 v1 2026-07-01T02:48:01.528Z