Related papers: Exponential families from a single KL identity
The matrix exponential restricted to skew-symmetric matrices has numerous applications, notably in view of its interpretation as the Lie group exponential and Riemannian exponential for the special orthogonal group. We characterize the…
We consider three classes of linear differential equations on distribution functions, with a fractional order $\alpha\in [0,1].$ The integer case $\alpha =1$ corresponds to the three classical extreme families. In general, we show that…
This document describes concisely the ubiquitous class of exponential family distributions met in statistics. The first part recalls definitions and summarizes main properties and duality with Bregman divergences (all proofs are skipped).…
The success of machine learning methods heavily relies on having an appropriate representation for data at hand. Traditionally, machine learning approaches relied on user-defined heuristics to extract features encoding structural…
In a regular full exponential family, the maximum likelihood estimator (MLE) need not exist in the traditional sense. However, the MLE may exist in the completion of the exponential family. Existing algorithms for finding the MLE in the…
The Kullback--Leibler divergence together with exponential families establishes the foundation of information geometry and is widely generalized. Among the generalization, we focus on the $(h,\tau)$-divergence and $(h,\tau)$-exponential…
Statistical inference may follow a frequentist approach or it may follow a Bayesian approach or it may use the minimum description length principle (MDL). Our goal is to identify situations in which these different approaches to statistical…
The recently successful Munchausen Reinforcement Learning (M-RL) features implicit Kullback-Leibler (KL) regularization by augmenting the reward function with logarithm of the current stochastic policy. Though significant improvement has…
The standard Large Deviation Theory (LDT) is mathematically illustrated by the Boltzmann-Gibbs factor which describes the thermal equilibrium of short-range-interacting many-body Hamiltonian systems, the velocity distribution of which is…
Modern variational inference (VI) uses stochastic gradients to avoid intractable expectations, enabling large-scale probabilistic inference in complex models. VI posits a family of approximating distributions q and then finds the member of…
There is accumulating evidence in the literature that stability of learning algorithms is a key characteristic that permits a learning algorithm to generalize. Despite various insightful results in this direction, there seems to be an…
We present an efficient algorithm for maximum likelihood estimation (MLE) of exponential family models, with a general parametrization of the energy function that includes neural networks. We exploit the primal-dual view of the MLE with a…
We generalize the exponential family of probability distributions. In our approach, the exponential function is replaced by a $\varphi$-function, resulting in a $\varphi$-family of probability distributions. We show how $\varphi$-families…
The property of learning-curve monotonicity, highlighted in a recent series of work by Loog, Mey and Viering, describes algorithms which only improve in average performance given more data, for any underlying data distribution within a…
Optimising probabilistic models is a well-studied field in statistics. However, its connection with the training of generative models remains largely under-explored. In this paper, we show that the evolution of time-varying generative…
We develop an analog of the exponential families of Wilf in which the label sets are finite dimensional vector spaces over a finite field rather than finite sets of positive integers. The essential features of exponential families are…
Using the technique developed in approximation theory, we construct examples of exponential families of infinitely divisible laws which can be viewed as deformations of the normal, gamma, and Poisson exponential families. Replacing the…
Exponential family extensions of principal component analysis (EPCA) have received a considerable amount of attention in recent years, demonstrating the growing need for basic modeling tools that do not assume the squared loss or Gaussian…
We consider the logarithm of the central value $\log L(1/2)$ in the orthogonal family ${L(s,f)}_{f \in H_k}$ where $H_k$ is the set of weight $k$ Hecke-eigen cusp form for $SL_2(\mathbb{Z})$, and in the symplectic family…
In numerous instances, the generalized exponential distribution can be used as an alternative to the most widely used non-regular family of distributions: Weibull, gamma, lognormal with three-parameters when analyzing lifetime or any skewed…