English
Related papers

Related papers: Measure Transport with Kernel Stein Discrepancy

200 papers

Persistence diagrams (PDs) play a key role in topological data analysis (TDA), in which they are routinely used to describe topological properties of complicated shapes. PDs enjoy strong stability properties and have proven their utility in…

Computational Geometry · Computer Science 2017-11-10 Mathieu Carrière , Marco Cuturi , Steve Oudot

Clustering is a fundamental unsupervised learning approach. Many clustering algorithms -- such as $k$-means -- rely on the euclidean distance as a similarity measure, which is often not the most relevant metric for high dimensional data…

Machine Learning · Computer Science 2019-10-22 Aude Genevay , Gabriel Dulac-Arnold , Jean-Philippe Vert

In this paper, we delve deeper into the Kullback-Leibler (KL) Divergence loss and mathematically prove that it is equivalent to the Decoupled Kullback-Leibler (DKL) Divergence loss that consists of 1) a weighted Mean Square Error (wMSE)…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Jiequan Cui , Zhuotao Tian , Zhisheng Zhong , Xiaojuan Qi , Bei Yu , Hanwang Zhang

Stein Variational Gradient Descent (SVGD) is a popular sampling algorithm used in various machine learning tasks. It is well known that SVGD arises from a discretization of the kernelized gradient flow of the Kullback-Leibler divergence…

Machine Learning · Computer Science 2022-11-22 Lukang Sun , Peter Richtárik

Sliced Stein discrepancy (SSD) and its kernelized variants have demonstrated promising successes in goodness-of-fit tests and model learning in high dimensions. Despite their theoretical elegance, their empirical performance depends…

Machine Learning · Computer Science 2021-07-22 Wenbo Gong , Kaibo Zhang , Yingzhen Li , José Miguel Hernández-Lobato

Representing, comparing, and measuring the distance between probability distributions is a key task in computational statistics and machine learning. The choice of representation and the associated distance determine properties of the…

Machine Learning · Statistics 2026-02-26 Masha Naslidnyk

We propose a general purpose variational inference algorithm that forms a natural counterpart of gradient descent for optimization. Our method iteratively transports a set of particles to match the target distribution, by applying a form of…

Machine Learning · Statistics 2019-09-10 Qiang Liu , Dilin Wang

In many contexts Gaussian Mixtures (GM) are used to approximate probability distributions, possibly time-varying. In some applications the number of GM components exponentially increases over time, and reduction procedures are required to…

Machine Learning · Statistics 2021-04-27 A. D'Ortenzio , C. Manes

The improvement in the performance of efficient and lightweight models (i.e., the student model) is achieved through knowledge distillation (KD), which involves transferring knowledge from more complex models (i.e., the teacher model).…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Seonghak Kim , Gyeongdo Ham , Yucheol Cho , Daeshik Kim

Machine learning (ML) has shown significant promise in studying complex geophysical dynamical systems, including turbulence and climate processes. Such systems often display sensitive dependence on initial conditions, reflected in positive…

Atmospheric and Oceanic Physics · Physics 2025-12-09 Zhewen Hou , Jiajin Sun , Subashree Venkatasubramanian , Peter Jin , Shuolin Li , Tian Zheng

Knowledge distillation (KD), transferring knowledge from a cumbersome teacher model to a lightweight student model, has been investigated to design efficient neural architectures. Generally, the objective function of KD is the…

Machine Learning · Computer Science 2021-05-20 Taehyeon Kim , Jaehoon Oh , NakYil Kim , Sangwook Cho , Se-Young Yun

We discuss a relation between the Kantorovich-Wasserstein (KW) metric and the Kullback-Leibler (KL) divergence. The former is defined using the optimal transport problem (OTP) in the Kantorovich formulation. The latter is used to define…

Information Theory · Computer Science 2019-08-27 Roman V. Belavkin

Embedding probability distributions into reproducing kernel Hilbert spaces (RKHS) has enabled powerful nonparametric methods such as the maximum mean discrepancy (MMD), a statistical distance with strong theoretical and computational…

Machine Learning · Statistics 2025-05-28 Masha Naslidnyk , Siu Lun Chau , François-Xavier Briol , Krikamol Muandet

Optimal transport is a geometrically intuitive, robust and flexible metric for sample comparison in data analysis and machine learning. Its formal Riemannian structure allows for a local linearization via a tangent space approximation. This…

Optimization and Control · Mathematics 2024-06-07 Clément Sarrazin , Bernhard Schmitzer

We examine the estimation of the Kullback-Leibler (KL) divergence and the use of the goodness-of-fit test for multivariate continuous distributions. Our starting point is the maximum entropy principle for Shannon entropy: among all…

Statistics Theory · Mathematics 2026-03-10 Mehmet Siddik Cadirci , Martin Singull

We consider the optimal mass transportation problem in $\RR^d$ with measurably parameterized marginals, for general cost functions and under conditions ensuring the existence of a unique optimal transport map. We prove a joint measurability…

Probability · Mathematics 2008-09-09 Joaquin Fontbona , Helene Guerin , Sylvie Meleard

Variational inference often minimizes the "reverse" Kullbeck-Leibler (KL) KL(q||p) from the approximate distribution q to the posterior p. Recent work studies the "forward" KL KL(p||q), which unlike reverse KL does not lead to variational…

Machine Learning · Statistics 2022-09-07 Liyi Zhang , David M. Blei , Christian A. Naesseth

We consider the problem of estimating probability density functions based on sample data, using a finite mixture of densities from some component class. To this end, we introduce the $h$-lifted Kullback--Leibler (KL) divergence as a…

Machine Learning · Statistics 2024-12-24 Mark Chiu Chong , Hien Duy Nguyen , TrungTin Nguyen

We derive a class of divergences measuring the difference between probability density functions on the one-dimensional sample space. This divergence is a one-parameter variation of the Itakura--Saito divergence between quantile density…

Information Theory · Computer Science 2026-02-20 Wuchen Li

This paper introduces kdiff, a novel kernel-based measure for estimating distances between instances of time series, random fields and other forms of structured data. This measure is based on the idea of matching distributions that only…

Machine Learning · Statistics 2021-10-01 Srinjoy Das , Hrushikesh Mhaskar , Alexander Cloninger
‹ Prev 1 3 4 5 6 7 10 Next ›