English
Related papers

Related papers: Alpha-Divergences in Variational Dropout

200 papers

The Gumbel-softmax distribution, or Concrete distribution, is often used to relax the discrete characteristics of a categorical distribution and enable back-propagation through differentiable reparameterization. Although it reliably yields…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-10 Sangshin Oh , Seyun Um , Hong-Goo Kang

Domain adaptation is an important problem and often needed for real-world applications. In this problem, instead of i.i.d. training and testing datapoints, we assume that the source (training) data and the target (testing) data have…

Machine Learning · Computer Science 2022-03-15 A. Tuan Nguyen , Toan Tran , Yarin Gal , Philip H. S. Torr , Atılım Güneş Baydin

We introduce a new family of particle evolution samplers suitable for constrained domains and non-Euclidean geometries. Stein Variational Mirror Descent and Mirrored Stein Variational Gradient Descent minimize the Kullback-Leibler (KL)…

Machine Learning · Statistics 2022-04-26 Jiaxin Shi , Chang Liu , Lester Mackey

The Variational Bayesian method (VB) is used to solve the probability distributions of latent variables with the minimum free energy criterion. This criterion is not easy to understand, and the computation is complex. For these reasons,…

Machine Learning · Computer Science 2026-05-01 Chenguang Lu

We consider the problem of time series forecasting in an adaptive setting. We focus on the inference of state-space models under unknown and potentially time-varying noise variances. We introduce an augmented model in which the variances…

Machine Learning · Computer Science 2021-11-10 Joseph de Vilmarest , Olivier Wintenberger

We introduce Group Spike-and-slab Variational Bayes (GSVB), a scalable method for group sparse regression. A fast co-ordinate ascent variational inference (CAVI) algorithm is developed for several common model families including Gaussian,…

Methodology · Statistics 2025-11-14 Michael Komodromos , Marina Evangelou , Sarah Filippi , Kolyan Ray

This article considers Bayesian model selection via mean-field (MF) variational approximation. Towards this goal, we study the non-asymptotic properties of MF inference under the Bayesian framework that allows latent variables and model…

Methodology · Statistics 2023-12-29 Yangfan Zhang , Yun Yang

I present an analytic method for estimating the errors in fitting a distribution. A well-known theorem from statistics gives the minimum variance bound (MVB) for the uncertainty in estimating a set of parameters $\l_i$, when a distribution…

Astrophysics · Physics 2009-10-22 Andrew Gould

Recently in [1, 2], Ali-Akbar Bromideh introduced the Kullback-Leibler Divergence (KLD) test statistic in discrim- inating between two models. It was found that the Ratio Minimized Kulback-Leibler Divergence (RMKLD) works better than the…

Methodology · Statistics 2017-10-02 Papa Ngom , Jean de Dieu Nkurunziza , Carlos Simplice Ogouyandjou

Mutual Information (MI) is a fundamental measure of statistical dependence widely used in representation learning. While direct optimization of MI via its definition as a Kullback-Leibler divergence (KLD) is often intractable, many recent…

Machine Learning · Computer Science 2026-03-18 Reuben Dorent , Polina Golland , William Wells

The Kullback-Leibler divergence, the Kullback-Leibler variation, and the Bernstein "norm" are used to quantify discrepancies among probability distributions in likelihood models such as nonparametric maximum likelihood and nonparametric…

Statistics Theory · Mathematics 2026-01-27 Tetsuya Kaji

In this work, we study an optimizer, Grad-Avg to optimize error functions. We establish the convergence of the sequence of iterates of Grad-Avg mathematically to a minimizer (under boundedness assumption). We apply Grad-Avg along with some…

Machine Learning · Computer Science 2020-12-11 Saugata Purkayastha , Sukannya Purkayastha

We study the fundamental and timely problem of learning long sequences in autoregressive modeling and next-token prediction under model misspecification, measured by the joint Kullback--Leibler (KL) divergence. Our goal is to characterize…

Machine Learning · Computer Science 2026-05-13 Yunbei Xu , Yuzhe Yuan , Ruohan Zhan

Approximating a probability density in a tractable manner is a central task in Bayesian statistics. Variational Inference (VI) is a popular technique that achieves tractability by choosing a relatively simple variational family. Borrowing…

Machine Learning · Statistics 2018-11-30 Francesco Locatello , Gideon Dresdner , Rajiv Khanna , Isabel Valera , Gunnar Rätsch

This work presents an upper-bound to value that the Kullback-Leibler (KL) divergence can reach for a class of probability distributions called quantum distributions (QD). The aim is to find a distribution $U$ which maximizes the KL…

Machine Learning · Computer Science 2020-12-11 Vincenzo Bonnici

We analyse the properties of an unbiased gradient estimator of the ELBO for variational inference, based on the score function method with leave-one-out control variates. We show that this gradient estimator can be obtained using a new…

Machine Learning · Statistics 2020-10-30 Lorenz Richter , Ayman Boustati , Nikolas Nüsken , Francisco J. R. Ruiz , Ömer Deniz Akyildiz

We derive a new variational formula for the R\'enyi family of divergences, $R_\alpha(Q\|P)$, between probability measures $Q$ and $P$. Our result generalizes the classical Donsker-Varadhan variational formula for the Kullback-Leibler…

Machine Learning · Statistics 2021-07-21 Jeremiah Birrell , Paul Dupuis , Markos A. Katsoulakis , Luc Rey-Bellet , Jie Wang

The recent boom in the literature on entropy-regularized reinforcement learning (RL) approaches reveals that Kullback-Leibler (KL) regularization brings advantages to RL algorithms by canceling out errors under mild assumptions. However,…

Machine Learning · Computer Science 2021-10-06 Toshinori Kitamura , Lingwei Zhu , Takamitsu Matsubara

We study the gradient flow for a relaxed approximation to the Kullback-Leibler (KL) divergence between a moving source and a fixed target distribution. This approximation, termed the KALE (KL approximate lower-bound estimator), solves a…

Machine Learning · Statistics 2021-11-01 Pierre Glaser , Michael Arbel , Arthur Gretton

The ability to compute the exact divergence between two high-dimensional distributions is useful in many applications but doing so naively is intractable. Computing the alpha-beta divergence -- a family of divergences that includes the…

Machine Learning · Computer Science 2023-10-17 Loong Kuan Lee , Geoffrey I. Webb , Daniel F. Schmidt , Nico Piatkowski
‹ Prev 1 8 9 10 Next ›