English

Towards explainable classifiers using the counterfactual approach -- global explanations for discovering bias in data

Machine Learning 2020-10-26 v2 Artificial Intelligence Computer Vision and Pattern Recognition Machine Learning

Abstract

The paper proposes summarized attribution-based post-hoc explanations for the detection and identification of bias in data. A global explanation is proposed, and a step-by-step framework on how to detect and test bias is introduced. Since removing unwanted bias is often a complicated and tremendous task, it is automatically inserted, instead. Then, the bias is evaluated with the proposed counterfactual approach. The obtained results are validated on a sample skin lesion dataset. Using the proposed method, a number of possible bias causing artifacts are successfully identified and confirmed in dermoscopy images. In particular, it is confirmed that black frames have a strong influence on Convolutional Neural Network's prediction: 22% of them changed the prediction from benign to malignant.

Keywords

Cite

@article{arxiv.2005.02269,
  title  = {Towards explainable classifiers using the counterfactual approach -- global explanations for discovering bias in data},
  author = {Agnieszka Mikołajczyk and Michał Grochowski and Arkadiusz Kwasigroch},
  journal= {arXiv preprint arXiv:2005.02269},
  year   = {2020}
}

Comments

Accepted for publication in Journal of Artificial Intelligence and Soft Computing Research; 12 pages, 4 figures, code available, 8-pages appendix

R2 v1 2026-06-23T15:19:37.918Z