English

A comparative analysis of machine learning algorithms for predicting probabilities of default

Risk Management 2025-06-25 v1 Machine Learning Applications

Abstract

Predicting the probability of default (PD) of prospective loans is a critical objective for financial institutions. In recent years, machine learning (ML) algorithms have achieved remarkable success across a wide variety of prediction tasks; yet, they remain relatively underutilised in credit risk analysis. This paper highlights the opportunities that ML algorithms offer to this field by comparing the performance of five predictive models-Random Forests, Decision Trees, XGBoost, Gradient Boosting and AdaBoost-to the predominantly used logistic regression, over a benchmark dataset from Scheule et al. (Credit Risk Analytics: The R Companion). Our findings underscore the strengths and weaknesses of each method, providing valuable insights into the most effective ML algorithms for PD prediction in the context of loan portfolios.

Keywords

Cite

@article{arxiv.2506.19789,
  title  = {A comparative analysis of machine learning algorithms for predicting probabilities of default},
  author = {Adrian Iulian Cristescu and Matteo Giordano},
  journal= {arXiv preprint arXiv:2506.19789},
  year   = {2025}
}

Comments

6 pages, 2 tables, to appear in Book of Short Papers - IES 2025

R2 v1 2026-07-01T03:31:54.708Z