English
Related papers

Related papers: Exponential families from a single KL identity

200 papers

We present the first algorithm for computing class groups and unit groups of arbitrary number fields that provably runs in probabilistic subexponential time, assuming the Extended Riemann Hypothesis (ERH). Previous subexponential algorithms…

Number Theory · Mathematics 2026-02-20 Koen de Boer , Alice Pellet-Mary , Benjamin Wesolowski

In this paper we study a class of exponential family on permutations, which includes some of the commonly studied Mallows models. We show that the pseudo-likelihood estimator for the natural parameter in the exponential family is…

Statistics Theory · Mathematics 2023-04-11 Sumit Mukherjee , Daiki Tagami

The prevailing assumption of an exponential decay in large language model (LLM) reliability with sequence length, predicated on independent per-token error probabilities, posits an inherent limitation for long autoregressive outputs. Our…

Computation and Language · Computer Science 2026-05-07 Mikhail L. Arbuzov , Sisong Bei , Ziwei Dong , Dmitri Kalaev , Alexey A. Shvets

We consider the challenging problem of statistical inference for exponential-family random graph models based on a single observation of a random graph with complex dependence. To facilitate statistical inference, we consider random graphs…

Statistics Theory · Mathematics 2020-03-13 Michael Schweinberger

Minimum divergence procedures based on the density power divergence and the logarithmic density power divergence have been extremely popular and successful in generating inference procedures which combine a high degree of model efficiency…

Statistics Theory · Mathematics 2022-11-10 Souvik Ray , Subrata Pal , Sumit Kumar Kar , Ayanendranath Basu

We study online learning under logarithmic loss with regular parametric models. Hedayati and Bartlett (2012b) showed that a Bayesian prediction strategy with Jeffreys prior and sequential normalized maximum likelihood (SNML) coincide and…

Machine Learning · Computer Science 2013-05-21 Peter Bartlett , Peter Grunwald , Peter Harremoes , Fares Hedayati , Wojciech Kotlowski

Consider a pair of random vectors $(\mathbf{X},\mathbf{Y}) $ and the conditional expectation operator $\mathbb{E}[\mathbf{X}|\mathbf{Y}=\mathbf{y}]$. This work studies analytic properties of the conditional expectation by characterizing…

Probability · Mathematics 2021-08-31 Alex Dytso , Martina Cardone

In this note we prove the dual representation formula of the divergence between two distributions in a parametric model. Resulting estimators for the divergence as for the parameter are derived. These estimators do not make use of any…

Methodology · Statistics 2011-08-23 Michel Broniatowski

The recently proposed Thermodynamic Variational Objective (TVO) leverages thermodynamic integration to provide a family of variational inference objectives, which both tighten and generalize the ubiquitous Evidence Lower Bound (ELBO).…

Machine Learning · Computer Science 2020-07-02 Rob Brekelmans , Vaden Masrani , Frank Wood , Greg Ver Steeg , Aram Galstyan

The link with exponential families has allowed $k$-means clustering to be generalized to a wide variety of data generating distributions in exponential families and clustering distortions among Bregman divergences. Getting the framework to…

Machine Learning · Computer Science 2022-11-08 Ehsan Amid , Richard Nock , Manfred Warmuth

Recent progress in center-based clustering algorithms combats poor local minima by implicit annealing, using a family of generalized means. These methods are variations of Lloyd's celebrated $k$-means algorithm, and are most appropriate for…

Machine Learning · Statistics 2022-06-23 Adithya Vellal , Saptarshi Chakraborty , Jason Xu

The ability of many powerful machine learning algorithms to deal with large data sets without compromise is often hampered by computationally expensive linear algebra tasks, of which calculating the log determinant is a canonical example.…

Machine Learning · Statistics 2017-09-11 Diego Granziol , Stephen Roberts

The generative nature of Large Language Models (LLMs) is reflected in the conditional probabilities they compute to sample each response token given the previous tokens. These probabilities encode the distributional structure that the model…

Computation and Language · Computer Science 2026-05-22 Shilpika Shilpika , Carlo Graziani , Bethany Lusch , Venkatram Vishwanath , Michael E. Papka

Limits of densities belonging to an exponential family appear in many applications, {e.g.} Gibbs models in Statistical Physics, relaxed combinatorial optimization, coding theory, critical likelihood computations, Bayes priors with singular…

Statistics Theory · Mathematics 2010-12-06 Luigi Malagò , Giovanni Pistone

While investigating the generalization of the Chandrasekhar (1943) dynamical friction to the case of field stars with a power-law mass spectrum and equipartition Maxwell-Boltzmann velocity distribution, a pair of 2-dimensional integrals…

Mathematical Physics · Physics 2020-09-15 Luca Ciotti

Free exponential families have been previously introduced as a special case of the q-exponential family. We show that free exponential families arise also from a procedure analogous to the definition of exponential families by using the…

Probability · Mathematics 2010-06-08 Wlodzimierz Bryc

Diffusion models are a new class of generative models that revolve around the estimation of the score function associated with a stochastic differential equation. Subsequent to its acquisition, the approximated score function is then…

Statistics Theory · Mathematics 2024-09-13 Giovanni Conforti , Alain Durmus , Marta Gentiloni Silveri

We address the problem of learning of continuous exponential family distributions with unbounded support. While a lot of progress has been made on learning of Gaussian graphical models, we still lack scalable algorithms for reconstructing…

Machine Learning · Computer Science 2022-03-01 Christopher X. Ren , Sidhant Misra , Marc Vuffray , Andrey Y. Lokhov

We derive a new variational formula for the R\'enyi family of divergences, $R_\alpha(Q\|P)$, between probability measures $Q$ and $P$. Our result generalizes the classical Donsker-Varadhan variational formula for the Kullback-Leibler…

Machine Learning · Statistics 2021-07-21 Jeremiah Birrell , Paul Dupuis , Markos A. Katsoulakis , Luc Rey-Bellet , Jie Wang

We prove that the exponential distribution is the only one which satisfies a regression identity. This identity involves conditional expectation of the sample mean of record values given two record values outside of the sample.

Probability · Mathematics 2011-05-06 George P. Yanev