English
Related papers

Related papers: Maximizing the Bregman divergence from a Bregman f…

200 papers

This paper investigates maximizers of the information divergence from an exponential family $E$. It is shown that the $rI$-projection of a maximizer $P$ to $E$ is a convex combination of $P$ and a probability measure $P_-$ with disjoint…

Information Theory · Computer Science 2014-06-18 Johannes Rauh

The bias-variance decomposition is a central result in statistics and machine learning, but is typically presented only for the squared error. We present a generalization of the bias-variance decomposition where the prediction error is a…

Machine Learning · Computer Science 2025-11-13 David Pfau

The {\lambda}-exponential family has recently been proposed to generalize the exponential family. While the exponential family is well-understood and widely used, this it not the case of the {\lambda}-exponential family. However, many…

Statistics Theory · Mathematics 2024-06-21 Thomas Guilmeau , Emilie Chouzenoux , Víctor Elvira

This article studies exponential families $\mathcal{E}$ on finite sets such that the information divergence $D(P\|\mathcal{E})$ of an arbitrary probability distribution from $\mathcal{E}$ is bounded by some constant $D>0$. A particular…

Statistics Theory · Mathematics 2014-06-18 Johannes Rauh

We introduce generalized notions of a divergence function and a Fisher information matrix. We propose to generalize the notion of an exponential family of models by reformulating it in terms of the Fisher information matrix. Our methods are…

Information Theory · Computer Science 2013-02-22 Jan Naudts , Ben Anthonis

Bregman divergences are a class of distance-like comparison functions which play fundamental roles in optimization, statistics, and information theory. One important property of Bregman divergences is that they cause two useful formulations…

Information Theory · Computer Science 2025-01-07 Philip S. Chodrow

We revisit the method of mixture technique, also known as the Laplace method, to study the concentration phenomenon in generic exponential families. Combining the properties of Bregman divergence associated with log-partition function of…

Machine Learning · Computer Science 2023-07-14 Sayak Ray Chowdhury , Patrick Saux , Odalric-Ambrym Maillard , Aditya Gopalan

We review recent results about the maximal values of the Kullback-Leibler information divergence from statistical models defined by neural networks, including naive Bayes models, restricted Boltzmann machines, deep belief networks, and…

Statistics Theory · Mathematics 2014-06-18 Guido Montufar , Johannes Rauh , Nihat Ay

Information divergence that measures the difference between two nonnegative matrices or tensors has found its use in a variety of machine learning problems. Examples are Nonnegative Matrix/Tensor Factorization, Stochastic Neighbor…

Machine Learning · Computer Science 2014-06-06 Onur Dikmen , Zhirong Yang , Erkki Oja

Exponential families form the backbone of modern statistics and machine learning, but textbooks seldom derive them from first principles in an accessible way. Although minimal sufficiency and the principle of maximum entropy, originating in…

Methodology · Statistics 2026-04-27 Korbinian Strimmer

This manuscript develops the theory of agglomerative clustering with Bregman divergences. Geometric smoothing techniques are developed to deal with degenerate clusters. To allow for cluster models based on exponential families with…

Machine Learning · Computer Science 2012-07-03 Matus Telgarsky , Sanjoy Dasgupta

Logarithmic score and information divergence appear in both information theory, statistics, statistical mechanics, and portfolio theory. We demonstrate that all these topics involve some kind of optimization that leads directly to the use…

Statistics Theory · Mathematics 2015-07-28 Peter Harremoës

We study the variational inference problem of minimizing a regularized R\'enyi divergence over an exponential family. We propose to solve this problem with a Bregman proximal gradient algorithm. We propose a sampling-based algorithm to…

Statistics Theory · Mathematics 2024-10-17 Thomas Guilmeau , Emilie Chouzenoux , Víctor Elvira

What if there is a teacher who knows the learning goal and wants to design good training data for a machine learner? We propose an optimal teaching framework aimed at learners who employ Bayesian models. Our framework is expressed as an…

Machine Learning · Computer Science 2013-10-04 Xiaojin Zhu

We introduce a new definition of exponential family of Markov chains, and show that many characteristic properties of the usual exponential family of probability distributions are properly extended to Markov chains. The method of…

Information Theory · Computer Science 2017-01-24 Hiroshi Nagaoka

The notion of Bregman divergence and sufficiency will be defined on general convex state spaces. It is demonstrated that only spectral sets can have a Bregman divergence that satisfies a sufficiency condition. Positive elements with trace 1…

Mathematical Physics · Physics 2017-07-17 Peter Harremoës

The paper introduces scaled Bregman distances of probability distributions which admit non-uniform contributions of observed events. They are introduced in a general form covering not only the distances of discrete and continuous stochastic…

Information Theory · Computer Science 2021-05-12 Wolfgang Stummer , Igor Vajda

Regularization by the Shannon entropy enables us to efficiently and approximately solve optimal transport problems on a finite set. This paper is concerned with regularized optimal transport problems via Bregman divergence. We introduce the…

Optimization and Control · Mathematics 2025-04-10 Keiichi Morikuni , Koya Sakakibara , Asuka Takatsu

A simple method is shown to provide optimal variational bounds on $f$-divergences with possible constraints on relative information extremums. Known results are refined or proved to be optimal as particular cases.

Information Theory · Computer Science 2019-02-05 Olivier Binette

Minimization of suitable statistical distances~(between the data and model densities) has proved to be a very useful technique in the field of robust inference. Apart from the class of $\phi$-divergences of \cite{a} and \cite{b}, the…

Statistics Theory · Mathematics 2021-01-25 Sancharee Basak , Ayanendranath Basu
‹ Prev 1 2 3 10 Next ›