English
Related papers

Related papers: Kullback-Leibler Divergence for the Normal-Gamma D…

200 papers

This work presents an infinite-dimensional generalization of the correspondence between the Kullback-Leibler and R\'enyi divergences between Gaussian measures on Euclidean space and the Alpha Log-Determinant divergences between symmetric,…

Probability · Mathematics 2019-04-12 Minh Ha Quang

Expectation maximization (EM) is the default algorithm for fitting probabilistic models with missing or latent variables, yet we lack a full understanding of its non-asymptotic convergence properties. Previous works show results along the…

Machine Learning · Computer Science 2022-03-01 Frederik Kunstner , Raunak Kumar , Mark Schmidt

We discuss Bayesian inference for a known-mean Gaussian model with a compound symmetric variance-covariance matrix. Since the space of such matrices is a linear subspace of that of positive definite matrices, we utilize the methods of…

Methodology · Statistics 2023-03-20 Zachary M. Pisano

Let $X| \mu \sim N_p(\mu,v_xI)$ and $Y| \mu \sim N_p(\mu,v_yI)$ be independent p-dimensional multivariate normal vectors with common unknown mean $\mu$. Based on only observing $X=x$, we consider the problem of obtaining a predictive…

Statistics Theory · Mathematics 2007-06-13 Edward I. George , Feng Liang , Xinyi Xu

In this work we introduce a family of transformations, named \textit{divergence transformations}, interpolating between any pair of probability density functions sharing the same support. We prove the remarkable property that the whole…

Mathematical Physics · Physics 2025-12-15 Razvan Gabriel Iagar , David Puertas-Centeno , Elio V. Toranzo

Simultaneous predictive densities for independent Poisson observables are investigated. The observed data and the target variables to be predicted are independently distributed according to different Poisson distributions parametrized by…

Statistics Theory · Mathematics 2021-05-27 Fumiyasu Komaki

To ensure stability of learning, state-of-the-art generalized policy iteration algorithms augment the policy improvement step with a trust region constraint bounding the information loss. The size of the trust region is commonly determined…

Machine Learning · Computer Science 2018-04-05 Boris Belousov , Jan Peters

Frequentist conditions for asymptotic suitability of Bayesian procedures focus on lower bounds for prior mass in Kullback-Leibler neighbourhoods of the data distribution. The goal of this paper is to investigate the flexibility in criteria…

Statistics Theory · Mathematics 2018-03-19 B. J. K. Kleijn , Y. Y. Zhao

We derive a deterministic, non-asymptotic upper bound on the Kullback-Leibler (KL) divergence of the flow-matching distribution approximation. In particular, if the $L_2$ flow-matching loss is bounded by $\epsilon^2 > 0$, then the KL…

Machine Learning · Computer Science 2025-11-10 Maojiang Su , Jerry Yao-Chieh Hu , Sophia Pi , Han Liu

The goal of this short note is to discuss the relation between Kullback--Leibler divergence and total variation distance, starting with the celebrated Pinsker's inequality relating the two, before switching to a simple, yet (arguably) more…

Probability · Mathematics 2023-08-03 Clément L. Canonne

In machine learning, it is common to optimize the parameters of a probabilistic model, modulated by an ad hoc regularization term that penalizes some values of the parameters. Regularization terms appear naturally in Variational Inference,…

Machine Learning · Computer Science 2024-02-08 Pierre Wolinski , Guillaume Charpiat , Yann Ollivier

In this study, simultaneous predictive distributions for independent Poisson observables were considered and the performance of predictive distributions was evaluated using the Kullback-Leibler (K-L) loss. This study proposes a class of…

Statistics Theory · Mathematics 2024-02-13 Xiao Li

This paper introduces link functions for transforming one probability distribution to another such that the Kullback-Leibler and R\'enyi divergences between the two distributions are symmetric. Two general classes of link models are…

Machine Learning · Statistics 2020-08-12 Majid Asadi , Karthik Devarajan , Nader Ebrahimi , Ehsan Soofi , Lauren Spirko-Burns

We provide a rigorous analysis of training by variational inference (VI) of Bayesian neural networks in the two-layer and infinite-width case. We consider a regression problem with a regularized evidence lower bound (ELBO) which is…

Machine Learning · Statistics 2023-07-12 Arnaud Descours , Tom Huix , Arnaud Guillin , Manon Michel , Éric Moulines , Boris Nectoux

We address the question of estimating Kullback-Leibler losses rather than squared losses in recovery problems where the noise is distributed within the exponential family. Inspired by Stein unbiased risk estimator (SURE), we exhibit…

Applications · Statistics 2017-08-22 Charles-Alban Deledalle

We study the minimax estimation of $\alpha$-divergences between discrete distributions for integer $\alpha\ge 1$, which include the Kullback--Leibler divergence and the $\chi^2$-divergences as special examples. Dropping the usual…

Information Theory · Computer Science 2021-03-04 Yanjun Han , Jiantao Jiao , Tsachy Weissman

Variational Inference (VI) is a popular alternative to asymptotically exact sampling in Bayesian inference. Its main workhorse is optimization over a reverse Kullback-Leibler divergence (RKL), which typically underestimates the tail of the…

Machine Learning · Statistics 2021-07-01 Ghassen Jerfel , Serena Wang , Clara Fannjiang , Katherine A. Heller , Yian Ma , Michael I. Jordan

Recent Reinforcement Learning (RL) algorithms making use of Kullback-Leibler (KL) regularization as a core component have shown outstanding performance. Yet, only little is understood theoretically about why KL regularization helps, so far.…

Machine Learning · Computer Science 2021-01-07 Nino Vieillard , Tadashi Kozuno , Bruno Scherrer , Olivier Pietquin , Rémi Munos , Matthieu Geist

Transfer learning, or domain adaptation, is concerned with machine learning problems in which training and testing data come from possibly different probability distributions. In this work, we give an information-theoretic analysis of the…

Information Theory · Computer Science 2024-08-09 Xuetong Wu , Jonathan H. Manton , Uwe Aickelin , Jingge Zhu

Bayesian coresets speed up posterior inference in the large-scale data regime by approximating the full-data log-likelihood function with a surrogate log-likelihood based on a small, weighted subset of the data. But while Bayesian coresets…

Machine Learning · Statistics 2024-10-18 Trevor Campbell