Related papers: Machine learning techniques in joint default asses…
In a world where Machine Learning (ML) is increasingly deployed to support decision-making in critical domains, providing decision-makers with explainable, stable, and relevant inputs becomes fundamental. Understanding how machine learning…
The stability of the financial system is associated with systemic risk factors such as the concurrent default of numerous small obligors. Hence it is of utmost importance to study the mutual dependence of losses for different creditors in…
This work introduces a novel framework for evaluating LLMs' capacity to balance instruction-following with critical reasoning when presented with multiple-choice questions containing no valid answers. Through systematic evaluation across…
In this work, we study the use of logistic regression in manufacturing failures detection. As a data set for the analysis, we used the data from Kaggle competition Bosch Production Line Performance. We considered the use of machine…
This paper generalizes Moody's correlated binomial default distribution for homogeneous (exchangeable) credit portfolio, which is introduced by Witt, to the case of inhomogeneous portfolios. As inhomogeneous portfolios, we consider two…
While defaults are rare events, losses can be substantial even for credit portfolios with a large number of contracts. Therefore, not only a good evaluation of the probability of default is crucial, but also the severity of losses needs to…
We discuss the parameter estimation of the probability of default (PD), the correlation between the obligors, and a phase transition. In our previous work, we studied the problem using the beta-binomial distribution. A non-equilibrium phase…
This paper presents a convenient framework for modeling default process and pricing derivative securities involving credit risk. The framework provides an integrated view of credit valuation adjustment by linking distance-to-default,…
Classification, the process of assigning a label (or class) to an observation given its features, is a common task in many applications. Nonetheless in most real-life applications, the labels can not be fully explained by the observed…
Non-stationary online learning has drawn much attention in recent years. Despite considerable progress, dynamic regret minimization has primarily focused on convex functions, leaving the functions with stronger curvature (e.g., squared or…
The application of deep learning to non-stationary temporal datasets can lead to overfitted models that underperform under regime changes. In this work, we propose a modular machine learning pipeline for ranking predictions on temporal…
Contrary to common belief, we show that gradient ascent-based unconstrained optimization methods frequently fail to perform machine unlearning, a phenomenon we attribute to the inherent statistical dependence between the forget and retain…
Purpose: Despite the potential of machine learning models, the lack of generalizability has hindered their widespread adoption in clinical practice. We investigate three methodological pitfalls: (1) violation of independence assumption, (2)…
Principal loading analysis is a dimension reduction method that discards variables which have only a small distorting effect on the covariance matrix. As a special case, principal loading analysis discards variables that are not correlated…
This paper presents PDx, an adaptive, machine learning operations (MLOps) driven decision system for forecasting credit risk using probability of default (PD) modeling in digital lending. While conventional PD models prioritize predictive…
Predictive models that are developed in a regulated industry or a regulated application, like determination of credit worthiness, must be interpretable and rational (e.g., meaningful improvements in basic credit behavior must result in…
Most machine learning techniques are based upon statistical learning theory, often simplified for the sake of computing speed. This paper is focused on the uncertainty aspect of mathematical modeling in machine learning. Regression analysis…
Multiple imputation is a common approach for dealing with missing values in statistical databases. The imputer fills in missing values with draws from predictive models estimated from the observed data, resulting in multiple, completed…
High-cardinality categorical variables are variables for which the number of different levels is large relative to the sample size of a data set, or in other words, there are few data points per level. Machine learning methods can have…
We compare observed corporate cumulative default probabilities to those calculated using a stochastic model based on an extension of the work of Black and Cox and find that corporations default as if via diffusive dynamics. The model, based…