Related papers: Exponential families from a single KL identity
We present the first algorithm for computing class groups and unit groups of arbitrary number fields that provably runs in probabilistic subexponential time, assuming the Extended Riemann Hypothesis (ERH). Previous subexponential algorithms…
In this paper we study a class of exponential family on permutations, which includes some of the commonly studied Mallows models. We show that the pseudo-likelihood estimator for the natural parameter in the exponential family is…
The prevailing assumption of an exponential decay in large language model (LLM) reliability with sequence length, predicated on independent per-token error probabilities, posits an inherent limitation for long autoregressive outputs. Our…
We consider the challenging problem of statistical inference for exponential-family random graph models based on a single observation of a random graph with complex dependence. To facilitate statistical inference, we consider random graphs…
Minimum divergence procedures based on the density power divergence and the logarithmic density power divergence have been extremely popular and successful in generating inference procedures which combine a high degree of model efficiency…
We study online learning under logarithmic loss with regular parametric models. Hedayati and Bartlett (2012b) showed that a Bayesian prediction strategy with Jeffreys prior and sequential normalized maximum likelihood (SNML) coincide and…
Consider a pair of random vectors $(\mathbf{X},\mathbf{Y}) $ and the conditional expectation operator $\mathbb{E}[\mathbf{X}|\mathbf{Y}=\mathbf{y}]$. This work studies analytic properties of the conditional expectation by characterizing…
In this note we prove the dual representation formula of the divergence between two distributions in a parametric model. Resulting estimators for the divergence as for the parameter are derived. These estimators do not make use of any…
The recently proposed Thermodynamic Variational Objective (TVO) leverages thermodynamic integration to provide a family of variational inference objectives, which both tighten and generalize the ubiquitous Evidence Lower Bound (ELBO).…
The link with exponential families has allowed $k$-means clustering to be generalized to a wide variety of data generating distributions in exponential families and clustering distortions among Bregman divergences. Getting the framework to…
Recent progress in center-based clustering algorithms combats poor local minima by implicit annealing, using a family of generalized means. These methods are variations of Lloyd's celebrated $k$-means algorithm, and are most appropriate for…
The ability of many powerful machine learning algorithms to deal with large data sets without compromise is often hampered by computationally expensive linear algebra tasks, of which calculating the log determinant is a canonical example.…
The generative nature of Large Language Models (LLMs) is reflected in the conditional probabilities they compute to sample each response token given the previous tokens. These probabilities encode the distributional structure that the model…
Limits of densities belonging to an exponential family appear in many applications, {e.g.} Gibbs models in Statistical Physics, relaxed combinatorial optimization, coding theory, critical likelihood computations, Bayes priors with singular…
While investigating the generalization of the Chandrasekhar (1943) dynamical friction to the case of field stars with a power-law mass spectrum and equipartition Maxwell-Boltzmann velocity distribution, a pair of 2-dimensional integrals…
Free exponential families have been previously introduced as a special case of the q-exponential family. We show that free exponential families arise also from a procedure analogous to the definition of exponential families by using the…
Diffusion models are a new class of generative models that revolve around the estimation of the score function associated with a stochastic differential equation. Subsequent to its acquisition, the approximated score function is then…
We address the problem of learning of continuous exponential family distributions with unbounded support. While a lot of progress has been made on learning of Gaussian graphical models, we still lack scalable algorithms for reconstructing…
We derive a new variational formula for the R\'enyi family of divergences, $R_\alpha(Q\|P)$, between probability measures $Q$ and $P$. Our result generalizes the classical Donsker-Varadhan variational formula for the Kullback-Leibler…
We prove that the exponential distribution is the only one which satisfies a regression identity. This identity involves conditional expectation of the sample mean of record values given two record values outside of the sample.