English
Related papers

Related papers: Relaxed Triangle Inequality for Kullback-Leibler D…

200 papers

We show that the $f$-divergence between any two densities of potentially different location-scale families can be reduced to the calculation of the $f$-divergence between one standard density with another location-scale density. It follows…

Statistics Theory · Mathematics 2021-02-16 Frank Nielsen

A loss function measures the discrepancy between the true values (observations) and their estimated fits, for a given instance of data. A loss function is said to be proper (unbiased, Fisher consistent) if the fits are defined over a unit…

Information Theory · Computer Science 2018-05-11 Amichai Painsky , Gregory W. Wornell

Symmetrized Kullback-Leibler (KL) information (\(I_{\mathrm{SKL}}\)), which symmetrizes the traditional mutual information by integrating Lautum information, has been shown as a critical quantity in communication~\cite{aminian2015capacity}…

Information Theory · Computer Science 2024-07-19 Haobo Chen , Gholamali Aminian , Yuheng Bu

In a regression setup with deterministic design, we study the pure aggregation problem and introduce a natural extension from the Gaussian distribution to distributions in the exponential family. While this extension bears strong…

Machine Learning · Statistics 2012-06-06 Philippe Rigollet

Knowing if a model will generalize to data 'in the wild' is crucial for safe deployment. To this end, we study model disagreement notions that consider the full predictive distribution - specifically disagreement based on Hellinger…

Machine Learning · Computer Science 2023-12-14 Mona Schirmer , Dan Zhang , Eric Nalisnick

The $\alpha$-divergences include the well-known Kullback-Leibler divergence, Hellinger distance and $\chi^2$-divergence. In this paper, we derive differential and integral relations between the $\alpha$-divergences that are generalizations…

Information Theory · Computer Science 2022-11-29 Tomohiro Nishiyama

Statistical divergences are ubiquitous in machine learning as tools for measuring discrepancy between probability distributions. As these applications inherently rely on approximating distributions from samples, we consider empirical…

Statistics Theory · Mathematics 2020-05-01 Ziv Goldfeld , Kengo Kato

The learning of Gaussian Mixture Models (also referred to simply as GMMs) plays an important role in machine learning. Known for their expressiveness and interpretability, Gaussian mixture models have a wide range of applications, from…

Machine Learning · Computer Science 2023-07-14 Ruichong Zhang

In this paper we prove the optimality of an aggregation procedure. We prove lower bounds for aggregation of model selection type of $M$ density estimators for the Kullback-Leiber divergence (KL), the Hellinger's distance and the…

Statistics Theory · Mathematics 2016-08-16 Guillaume Lecué

The families of $f$-divergences (e.g. the Kullback-Leibler divergence) and Integral Probability Metrics (e.g. total variation distance or maximum mean discrepancies) are widely used to quantify the similarity between probability…

Statistics Theory · Mathematics 2021-06-08 Rohit Agrawal , Thibaut Horel

This paper mainly focuses on studying the Shannon Entropy and Kullback-Leibler divergence of the multivariate log canonical fundamental skew-normal (LCFUSN) and canonical fundamental skew-normal (CFUSN) families of distributions, extending…

Statistics Theory · Mathematics 2016-09-21 Marina Muniz , Roger Silva , Rosangela Loschi

One of the most popular algorithms for clustering in Euclidean space is the $k$-means algorithm; $k$-means is difficult to analyze mathematically, and few theoretical guarantees are known about it, particularly when the data is {\em…

Machine Learning · Computer Science 2009-12-02 Kamalika Chaudhuri , Sanjoy Dasgupta , Andrea Vattani

The Poisson model is frequently employed to describe count data, but in a Bayesian context it leads to an analytically intractable posterior probability distribution. In this work, we analyze a variational Gaussian approximation to the…

Numerical Analysis · Mathematics 2018-02-14 Simon Arridge , Kazufumi Ito , Bangti Jin , Chen Zhang

We address the problem of detecting a change in the distribution of a high-dimensional multivariate normal time series. Assuming that the post-change parameters are unknown and estimated using a window of historical data, we extend the…

Signal Processing · Electrical Eng. & Systems 2025-02-12 Robert Malinas , Dogyoon Song , Benjamin D. Robinson , Alfred O. Hero

Orthogonal nonnegative matrix factorization (ONMF) has become a standard approach for clustering. As far as we know, most works on ONMF rely on the Frobenius norm to assess the quality of the approximation. This paper presents a new model…

Machine Learning · Statistics 2025-11-06 Jean Pacifique Nkurunziza , Fulgence Nahayo , Nicolas Gillis

The Jensen-Shannon divergence is a renown bounded symmetrization of the unbounded Kullback-Leibler divergence which measures the total Kullback-Leibler divergence to the average mixture distribution. However the Jensen-Shannon divergence…

Information Theory · Computer Science 2022-09-21 Frank Nielsen

The purpose of this paper is to examine the sampling problem through Euler discretization, where the potential function is assumed to be a mixture of locally smooth distributions and weakly dissipative. We introduce $\alpha_{G}$-mixture…

Computation · Statistics 2023-02-01 Dao Nguyen

This paper shows that large nonparametric classes of conditional multivariate densities can be approximated in the Kullback--Leibler distance by different specifications of finite mixtures of normal regressions in which normal means and…

Statistics Theory · Mathematics 2010-10-05 Andriy Norets

We consider the problem of sampling from a probability distribution $\pi$ which admits a density w.r.t. a dominating measure. It is well known that this can be written as an optimisation problem over the space of probability distributions…

Methodology · Statistics 2026-05-06 Francesca Romana Crucinio

It is well known and readily seen that the maximum of $n$ independent and uniformly on $[0,1]$ distributed random variables, suitably standardised, converges in total variation distance, as $n$ increases, to the standard negative…

Probability · Mathematics 2020-05-06 Michael Falk , Simone A. Padoan , Stefano Rizzelli