English
Related papers

Related papers: On the Universality of the Logistic Loss Function

200 papers

In statistical classification/multiple hypothesis testing and machine learning, a model distribution estimated from the training data is usually applied to replace the unknown true distribution in the Bayes decision rule, which introduces a…

Information Theory · Computer Science 2024-09-24 Zijian Yang , Vahe Eminyan , Ralf Schlüter , Hermann Ney

In a variety of applications it is important to extract information from a probability measure $\mu$ on an infinite dimensional space. Examples include the Bayesian approach to inverse problems and possibly conditioned) continuous time…

Probability · Mathematics 2016-06-02 Frank Pinski , Gideon Simpson , Andrew Stuart , Hendrik Weber

We study losses for binary classification and class probability estimation and extend the understanding of them from margin losses to general composite losses which are the composition of a proper loss with a link function. We characterise…

Machine Learning · Statistics 2009-12-18 Mark D. Reid , Robert C. Williamson

Coupling arguments are a central tool for bounding the deviation between two stochastic processes, but traditionally have been limited to Wasserstein metrics. In this paper, we apply the shifted composition rule--an information-theoretic…

Statistics Theory · Mathematics 2024-12-25 Jason M. Altschuler , Sinho Chewi

Estimating the Kullback-Leibler (KL) divergence between two distributions given samples from them is well-studied in machine learning and information theory. Motivated by considerations of multi-group fairness, we seek KL divergence…

Machine Learning · Computer Science 2022-03-01 Parikshit Gopalan , Nina Narodytska , Omer Reingold , Vatsal Sharan , Udi Wieder

Multi-dimensional distributions whose marginal distributions are uniform are called copulas. Among them, the one that satisfies given constraints on expectation and is closest to the independent distribution in the sense of Kullback-Leibler…

Methodology · Statistics 2022-04-11 Yici Chen , Tomonari Sei

Bayesian neural networks (BNNs) are state-of-the-art machine learning methods that can naturally regularize and systematically quantify uncertainties using their stochastic parameters. Kullback-Leibler (KL) divergence-based variational…

Machine Learning · Computer Science 2024-12-10 Ponkrshnan Thiagarajan , Susanta Ghosh

It has been argued persuasively that, in order to evaluate climate models, the probability distributions of model output need to be compared to the corresponding empirical distributions of observed data. Distance measures between…

Methodology · Statistics 2013-07-17 Thordis L. Thorarinsdottir , Tilmann Gneiting , Nadine Gissibl

We consider the estimation of the slope function in functional linear regression, where scalar responses are modeled in dependence of random functions. Cardot and Johannes [J. Multivariate Anal. 101 (2010) 395-408] have shown that a…

Statistics Theory · Mathematics 2013-02-19 Fabienne Comte , Jan Johannes

This paper illustrates the central role of loss functions in data-driven decision making, providing a comprehensive survey on their influence in cost-sensitive classification (CSC) and reinforcement learning (RL). We demonstrate how…

Machine Learning · Statistics 2025-04-07 Kaiwen Wang , Nathan Kallus , Wen Sun

Bernstein-von Mises results (BvM) establish that the Laplace approximation is asymptotically correct in the large-data limit. However, these results are inappropriate for computational purposes since they only hold over most, and not all,…

Statistics Theory · Mathematics 2019-05-01 Guillaume P. Dehaene

The forward Kullback-Leibler (KL) divergence is a ubiquitous objective for fitting a parameterized distribution to samples due to its tractability and equivalence to maximum likelihood estimation (MLE). Its inherent asymmetry, however, may…

Machine Learning · Computer Science 2026-05-12 Omri Ben-Dov , Luiz F. O. Chamon

This work introduces a general framework for calibeating based on regret minimization. As compared to Foster and Hart's seminal calibeating work which had specialized treatments of Brier score (squared loss) and log loss, we consider a…

Machine Learning · Computer Science 2026-05-19 Maximilian Fichtl , Cristóbal Guzmán , Nishant A. Mehta

Current approaches in approximate inference for Bayesian neural networks minimise the Kullback-Leibler divergence to approximate the true posterior over the weights. However, this approximation is without knowledge of the final application,…

Machine Learning · Statistics 2018-05-11 Adam D. Cobb , Stephen J. Roberts , Yarin Gal

The performance of machine learning classification algorithms are evaluated by estimating metrics, often from the confusion matrix, using training data and cross-validation. However, these do not prove that the best possible performance has…

Machine Learning · Statistics 2024-03-05 L. Crow , S. J. Watts

We propose a non-parametric variant of binary regression, where the hypothesis is regularized to be a Lipschitz function taking a metric space to [0,1] and the loss is logarithmic. This setting presents novel computational and statistical…

Machine Learning · Computer Science 2020-10-21 Ariel Avital , Klim Efremenko , Aryeh Kontorovich , David Toplin , Bo Waggoner

Mixability of a loss is known to characterise when constant regret bounds are achievable in games of prediction with expert advice through the use of Vovk's aggregating algorithm. We provide a new interpretation of mixability via convex…

Machine Learning · Computer Science 2014-03-12 Mark D. Reid , Rafael M. Frongillo , Robert C. Williamson

Kullback-Leibler divergence (KL) regularization is widely used in reinforcement learning, but it becomes infinite under support mismatch and can degenerate in low-noise limits. Utilizing a unified information-geometric framework, we…

Optimization and Control · Mathematics 2026-02-03 Viktor Stein , Adwait Datar , Nihat Ay

In this paper, we prove the universal consistency of wide and deep ReLU neural network classifiers trained on the logistic loss. We also give sufficient conditions for a class of probability measures for which classifiers based on neural…

Machine Learning · Statistics 2024-02-01 Hyunouk Ko , Xiaoming Huo

We examine the estimation of the Kullback-Leibler (KL) divergence and the use of the goodness-of-fit test for multivariate continuous distributions. Our starting point is the maximum entropy principle for Shannon entropy: among all…

Statistics Theory · Mathematics 2026-03-10 Mehmet Siddik Cadirci , Martin Singull