English
Related papers

Related papers: Moreau-Yosida $f$-divergences

200 papers

Bayesian Neural Networks (BNNs) are trained to optimize an entire distribution over their weights instead of a single set, having significant advantages in terms of, e.g., interpretability, multi-task learning, and calibration. Because of…

Machine Learning · Computer Science 2022-10-07 Jary Pomponi , Simone Scardapane , Aurelio Uncini

We propose a novel Riemannian geometric framework for variational inference in Bayesian models based on the nonparametric Fisher-Rao metric on the manifold of probability density functions. Under the square-root density representation, the…

Methodology · Statistics 2019-03-29 Abhijoy Saha , Karthik Bharath , Sebastian Kurtek

This paper introduces a variational approximation framework using direct optimization of what is known as the {\it scale invariant Alpha-Beta divergence} (sAB divergence). This new objective encompasses most variational objectives that use…

Machine Learning · Statistics 2018-05-22 Jean-Baptiste Regli , Ricardo Silva

Generative adversarial networks (GANs) have enjoyed much success in learning high-dimensional distributions. Learning objectives approximately minimize an $f$-divergence ($f$-GANs) or an integral probability metric (Wasserstein GANs)…

Machine Learning · Computer Science 2020-06-19 Jiaming Song , Stefano Ermon

This paper develops systematic approaches to obtain $f$-divergence inequalities, dealing with pairs of probability measures defined on arbitrary alphabets. Functional domination is one such approach, where special emphasis is placed on…

Information Theory · Computer Science 2016-12-06 Igal Sason , Sergio Verdú

Probabilistic models are often trained by maximum likelihood, which corresponds to minimizing a specific f-divergence between the model and data distribution. In light of recent successes in training Generative Adversarial Networks,…

Machine Learning · Statistics 2024-12-17 Mingtian Zhang , Thomas Bird , Raza Habib , Tianlin Xu , David Barber

Mutual Information (MI) is a fundamental measure of statistical dependence widely used in representation learning. While direct optimization of MI via its definition as a Kullback-Leibler divergence (KLD) is often intractable, many recent…

Machine Learning · Computer Science 2026-03-18 Reuben Dorent , Polina Golland , William Wells

In this paper, we provide three applications for $f$-divergences: (i) we introduce Sanov's upper bound on the tail probability of the sum of independent random variables based on super-modular $f$-divergence and show that our generalized…

Information Theory · Computer Science 2023-01-27 Saeed Masiha , Amin Gohari , Mohammad Hossein Yassaee

The article is devoted to the development of numerical methods for solving variational inequalities with relatively strongly monotone operators. We consider two classes of variational inequalities related to some analogs of the Lipschitz…

Optimization and Control · Mathematics 2022-05-25 F. S. Stonyakin , A. A. Titov , D. V. Makarenko , M. S. Alkousa

From a variational perspective, many statistical learning criteria involve seeking a distribution that balances empirical risk and regularization. In this paper, we broaden this perspective by introducing a new general class of variational…

Machine Learning · Computer Science 2026-02-17 Sophia Sklaviadis , Thomas Moellenhoff , Andre Martins , Mario Figueiredo

This paper introduces the $f$-divergence variational inference ($f$-VI) that generalizes variational inference to all $f$-divergences. Initiated from minimizing a crafty surrogate $f$-divergence that shares the statistical consistency with…

Machine Learning · Computer Science 2021-04-06 Neng Wan , Dapeng Li , Naira Hovakimyan

For any given neural network architecture a permutation of weights and biases results in the same functional network. This implies that optimization algorithms used to `train' or `learn' the network are faced with a very large number (in…

Optimization and Control · Mathematics 2022-02-22 Harbir Antil , Thomas S. Brown , Rainald Löhner , Fumiya Togashi , Deepanshu Verma

Generative Adversarial Networks (GANs) have been widely adopted in various fields. However, existing GANs generally are not able to preserve the manifold of data space, mainly due to the simple representation of discriminator for the…

Computer Vision and Pattern Recognition · Computer Science 2021-09-21 Haozhe Liu , Hanbang Liang , Xianxu Hou , Haoqian Wu , Feng Liu , Linlin Shen

Finite differences, as a subclass of direct methods in the calculus of variations, consist in discretizing the objective functional using appropriate approximations for derivatives that appear in the problem. This article generalizes the…

Optimization and Control · Mathematics 2013-08-09 Shakoor Pooseh , Ricardo Almeida , Delfim F. M. Torres

Fused Gromov-Wasserstein (FGW) distances provide a principled framework for comparing objects by jointly aligning structure and node features. However, existing FGW formulations treat all features uniformly, which limits interpretability…

Machine Learning · Computer Science 2026-05-13 Harlin Lee , Ying Yu , Mingxin Li , Ranthony Clark

Data processing inequalities for $f$-divergences can be sharpened using constants called "contraction coefficients" to produce strong data processing inequalities. For any discrete source-channel pair, the contraction coefficients for…

Information Theory · Computer Science 2018-07-17 Anuran Makur , Lizhong Zheng

This paper shows that large nonparametric classes of conditional multivariate densities can be approximated in the Kullback--Leibler distance by different specifications of finite mixtures of normal regressions in which normal means and…

Statistics Theory · Mathematics 2010-10-05 Andriy Norets

We present a novel approach to approximate Gaussian and mixture-of-Gaussians filtering. Our method relies on a variational approximation via a gradient-flow representation. The gradient flow is derived from a Kullback--Leibler discrepancy…

Computation · Statistics 2023-06-21 Adrien Corenflos , Hany Abdulsamad

Developing deep generative models that flexibly incorporate diverse measures of probability distance is an important area of research. Here we develop an unified mathematical framework of f-divergence generative model, f-GM, that…

Machine Learning · Statistics 2022-05-12 Jaime Roquero Gimenez , James Zou

Latent variable models are powerful tools for learning low-dimensional manifolds from high-dimensional data. However, when dealing with constrained data such as unit-norm vectors or symmetric positive-definite matrices, existing approaches…

Machine Learning · Computer Science 2025-03-10 Leonel Rozo , Miguel González-Duque , Noémie Jaquier , Søren Hauberg