English

Predicting credit default probabilities using machine learning techniques in the face of unequal class distributions

Econometrics 2019-07-31 v1 Machine Learning

Abstract

This study conducts a benchmarking study, comparing 23 different statistical and machine learning methods in a credit scoring application. In order to do so, the models' performance is evaluated over four different data sets in combination with five data sampling strategies to tackle existing class imbalances in the data. Six different performance measures are used to cover different aspects of predictive performance. The results indicate a strong superiority of ensemble methods and show that simple sampling strategies deliver better results than more sophisticated ones.

Keywords

Cite

@article{arxiv.1907.12996,
  title  = {Predicting credit default probabilities using machine learning techniques in the face of unequal class distributions},
  author = {Anna Stelzer},
  journal= {arXiv preprint arXiv:1907.12996},
  year   = {2019}
}

Comments

Keywords: Forecasting, credit scoring, imbalanced data sets, classification, benchmarking

R2 v1 2026-06-23T10:34:57.196Z