Related papers: On the Kullback-Leibler divergence between locatio…
Selecting an appropriate divergence measure is a critical aspect of machine learning, as it directly impacts model performance. Among the most widely used, we find the Kullback-Leibler (KL) divergence, originally introduced in kinetic…
We investigate random walks in independent, identically distributed random sceneries under the assumption that the scenery variables satisfy Cramer's condition. We prove moderate deviation principles in dimensions two and larger, covering…
We derive a class of divergences measuring the difference between probability density functions on the one-dimensional sample space. This divergence is a one-parameter variation of the Itakura--Saito divergence between quantile density…
The gapped local alignment score of two random sequences follows a Gumbel distribution. If computers could estimate the parameters of the Gumbel distribution within one second, the use of arbitrary alignment scoring schemes could increase…
We investigate the effects on cosmological clustering statistics of empirical biasing, where the galaxy distribution is a local transformation of the present-day Eulerian density field. The effects of the suppression of galaxy numbers in…
In this paper, we consider an infinite dimensional exponential family, $\mathcal{P}$ of probability densities, which are parametrized by functions in a reproducing kernel Hilbert space, $H$ and show it to be quite rich in the sense that a…
We derive a closed form solution for the Kullback-Leibler divergence between two Weibull distributions. These notes are meant as reference material and intended to provide a guided tour towards a result that is often mentioned but seldom…
Bernstein-von Mises results (BvM) establish that the Laplace approximation is asymptotically correct in the large-data limit. However, these results are inappropriate for computational purposes since they only hold over most, and not all,…
We study the maximum likelihood estimator of density of $n$ independent observations, under the assumption that it is well approximated by a mixture with a large number of components. The main focus is on statistical properties with respect…
We examine the geometry of the spaces between particles in diffusion-limited cluster aggregation, a numerical model of aggregating suspensions. Computing the distribution of distances from each point to the nearest particle, we show that it…
The performance of machine learning classification algorithms are evaluated by estimating metrics, often from the confusion matrix, using training data and cross-validation. However, these do not prove that the best possible performance has…
Variational Inference approximates an unnormalized distribution via the minimization of Kullback-Leibler (KL) divergence. Although this divergence is efficient for computation and has been widely used in applications, it suffers from some…
In this paper, we derive some upper and lower bounds and inequalities for the total variation distance (TVD) and the Kullback-Leibler divergence (KLD), also known as the relative entropy, between two probability measures $\mu$ and $\nu$…
Diffusion models have achieved great success in generating high-dimensional samples across various applications. While the theoretical guarantees for continuous-state diffusion models have been extensively studied, the convergence analysis…
We give a lower bound for the size of a subset of $\mathbb F_q^n$ containing a rich k-plane in every direction, a k-plane Furstenberg set. The chief novelty of our method is that we use arguments on non-reduced subschemes and flat families…
This article discusses the problem of estimation of parameters in finite mixtures when the mixture components are assumed to be symmetric and to come from the same location family. We refer to these mixtures as semi-parametric because no…
I outline the connections between some of the most widely used statistical measures of galaxy clustering and the fundamental issues in the theory of structure formation. I devote particular attention to the problem of biasing, i.e. to a…
We give an improved theoretical analysis of score-based generative modeling. Under a score estimate with small $L^2$ error (averaged across timesteps), we provide efficient convergence guarantees for any data distribution with second-order…
In this paper, we use the class of Wasserstein metrics to study asymptotic properties of posterior distributions. Our first goal is to provide sufficient conditions for posterior consistency. In addition to the well-known Schwartz's…
Exponential families are statistical models which are the workhorses in statistics, information theory, and machine learning among others. An exponential family can either be normalized subtractively by its cumulant or free energy function…