English
Related papers

Related papers: On the Universality of the Logistic Loss Function

200 papers

I introduce two novel loss functions for classification in deep learning. The two loss functions extend standard cross entropy loss by regularizing it with minimum entropy and Kullback-Leibler (K-L) divergence terms. The first of the two…

Computer Vision and Pattern Recognition · Computer Science 2025-01-24 Abdulrahman Oladipupo Ibraheem

We present a new perspective on the celebrated Sinkhorn algorithm by showing that is a special case of incremental/stochastic mirror descent. In order to see this, one should simply plug Kullback-Leibler divergence in both mirror map and…

Machine Learning · Computer Science 2019-09-17 Konstantin Mishchenko

We propose a dimension reduction technique for Bayesian inverse problems with nonlinear forward operators, non-Gaussian priors, and non-Gaussian observation noise. The likelihood function is approximated by a ridge function, i.e., a map…

Probability · Mathematics 2022-01-31 Olivier Zahm , Tiangang Cui , Kody Law , Alessio Spantini , Youssef Marzouk

This paper considers reparameterization invariant Bayesian point estimates and credible regions of model parameters for scientific inference and communication. The effect of intrinsic loss function choice in Bayesian intrinsic estimates and…

Methodology · Statistics 2021-09-23 Aki Vehtari

The loss function is arguably among the most important hyperparameters for a neural network. Many loss functions have been designed to date, making a correct choice nontrivial. However, elaborate justifications regarding the choice of the…

Machine Learning · Computer Science 2022-10-31 Simon Dräger , Jannik Dunkelau

It is widely conjectured that the reason that training algorithms for neural networks are successful because all local minima lead to similar performance, for example, see (LeCun et al., 2015, Choromanska et al., 2015, Dauphin et al.,…

Machine Learning · Computer Science 2018-03-06 Shiyu Liang , Ruoyu Sun , Yixuan Li , R. Srikant

We study the fundamental and timely problem of learning long sequences in autoregressive modeling and next-token prediction under model misspecification, measured by the joint Kullback--Leibler (KL) divergence. Our goal is to characterize…

Machine Learning · Computer Science 2026-05-13 Yunbei Xu , Yuzhe Yuan , Ruohan Zhan

Density-based directed distances -- particularly known as divergences -- between probability distributions are widely used in statistics as well as in the adjacent research fields of information theory, artificial intelligence and machine…

Statistics Theory · Mathematics 2022-03-03 Michel Broniatowski , Wolfgang Stummer

The maximum likelihood method is the best-known method for estimating the probabilities behind the data. However, the conventional method obtains the probability model closest to the empirical distribution, resulting in overfitting. Then…

Machine Learning · Statistics 2023-10-03 Akihisa Ichiki

In this paper we extend the setting of the online prediction with expert advice to function-valued forecasts. At each step of the online game several experts predict a function, and the learner has to efficiently aggregate these functional…

Machine Learning · Computer Science 2025-01-07 Alexander Korotin , Vladimir V'yugin , Evgeny Burnaev

This book deals with functions allowing to express the dissimilarity (discrepancy) between two data fields or ''divergence functions'' with the aim of applications to linear inverse problems. Most of the divergences found in the litterature…

Optimization and Control · Mathematics 2020-03-04 Henri Lantéri

We prove an L2 recovery bound for a family of sparse estimators defined as minimizers of some empirical loss functions -- which include hinge loss and logistic loss. More precisely, we achieve an upper-bound for coefficients estimation…

Statistics Theory · Mathematics 2019-01-15 Antoine Dedieu

Loss functions are at the heart of deep learning, shaping how models learn and perform across diverse tasks. They are used to quantify the difference between predicted outputs and ground truth labels, guiding the optimization process to…

Machine Learning · Computer Science 2025-09-11 Omar Elharrouss , Yasir Mahmood , Yassine Bechqito , Mohamed Adel Serhani , Elarbi Badidi , Jamal Riffi , Hamid Tairi

We develop new approaches in multi-class settings for constructing proper scoring rules and hinge-like losses and establishing corresponding regret bounds with respect to the zero-one or cost-weighted classification loss. Our construction…

Statistics Theory · Mathematics 2021-05-18 Zhiqiang Tan , Xinwei Zhang

A well-known technique in estimating probabilities of rare events in general and in information theory in particular (used, e.g., in the sphere-packing bound), is that of finding a reference probability measure under which the event of…

Information Theory · Computer Science 2014-12-23 Rami Atar , Neri Merhav

In this paper, we delve deeper into the Kullback-Leibler (KL) Divergence loss and mathematically prove that it is equivalent to the Decoupled Kullback-Leibler (DKL) Divergence loss that consists of 1) a weighted Mean Square Error (wMSE)…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Jiequan Cui , Zhuotao Tian , Zhisheng Zhong , Xiaojuan Qi , Bei Yu , Hanwang Zhang

We show that the log-likelihood of several probabilistic graphical models is Lipschitz continuous with respect to the lp-norm of the parameters. We discuss several implications of Lipschitz parametrization. We present an upper bound of the…

Machine Learning · Computer Science 2018-11-16 Jean Honorio

This paper develops a framework for fitting functions with domains in the Euclidean space, when data are sparse but a slow variation allows for a useful fit. We measure the variation by Lipschitz Bound (LB). Functions which admit smaller LB…

Methodology · Statistics 2014-07-07 Reza Hosseini , Akimichi Takemura , Kiros Berhane

The goal of binary classification is to estimate a discriminant function $\gamma$ from observations of covariate vectors and corresponding binary labels. We consider an elaboration of this problem in which the covariates are not available…

Statistics Theory · Mathematics 2009-09-29 XuanLong Nguyen , Martin J. Wainwright , Michael I. Jordan

Logarithmic score and information divergence appear in information theory, statistics, statistical mechanics, and portfolio theory. We demonstrate that all these topics involve some kind of optimization that leads directly to regret…

Information Theory · Computer Science 2017-07-17 Peter Harremoës
‹ Prev 1 3 4 5 6 7 10 Next ›