Related papers: Symmetric KL-divergence by Stein's Method
In this paper, we derive a unified method for establishing the distributional convergence of linear eigenvalue statistics (LES) for generalized patterned random matrices. We prove that for an $N \times N$ generalized patterned random matrix…
Gradient information on the sampling distribution can be used to reduce the variance of Monte Carlo estimators via Stein's method. An important application is that of estimating an expectation of a test function along the sample path of a…
We use a new method via $p$-Wasserstein bounds to prove Cram\'er-type moderate deviations in (multivariate) normal approximations. In the classical setting that $W$ is a standardized sum of $n$ independent and identically distributed…
Motivated by the central limit problem for convex bodies, we study normal approximation of linear functionals of high-dimensional random vectors with various types of symmetries. In particular, we obtain results for distributions which are…
Stein Variational Gradient Descent (SVGD) is a deterministic interacting-particle method for sampling from a target probability measure given access to its score function. In the mean-field and continuous-time limit, it is known that the…
We use Stein's method to obtain a bound on the distance between scaled $p$-dimensional random walks and a $p$-dimensional (correlated) Brownian Motion. We consider dependence schemes including those in which the summands in scaled sums are…
Gaussian processes are distributions over functions that are versatile and mathematically convenient priors in Bayesian modelling. However, their use is often impeded for data with large numbers of observations, $N$, due to the cubic (in…
We use Stein's method to bound the Wasserstein distance of order $2$ between a measure $\nu$ and the Gaussian measure using a stochastic process $(X_t)_{t \geq 0}$ such that $X_t$ is drawn from $\nu$ for any $t > 0$. If the stochastic…
We consider sequences of random variables of the type $S_n= n^{-1/2} \sum_{k=1}^n \{f(X_k)-\E[f(X_k)]\}$, $n\geq 1$, where $X=(X_k)_{k\in \Z}$ is a $d$-dimensional Gaussian process and $f: \R^d \rightarrow \R$ is a measurable function. It…
We analyze online and mini-batch k-means variants. Both scale up the widely used Lloyd 's algorithm via stochastic approximation, and have become popular for large-scale clustering and unsupervised feature learning. We show, for the first…
For a given positive random variable $V>0$ and a given $Z\sim N(0,1)$ independent of $V$, we compute the scalar $t_0$ such that the distance between $Z\sqrt{V}$ and $Z\sqrt{t_0}$ in the $L^2(\R)$ sense, is minimal. We also consider the same…
We examine the estimation of the Kullback-Leibler (KL) divergence and the use of the goodness-of-fit test for multivariate continuous distributions. Our starting point is the maximum entropy principle for Shannon entropy: among all…
On a probability space $(\Omega, \mathcal F, \mathbb P)$ we consider two independent sequences $(a_k)_{k \geq 1}$ and $(b_k)_{k \geq 1}$ of i.i.d. random variables that are centered with unit variance and which admit a moment strictly…
We give an improved theoretical analysis of score-based generative modeling. Under a score estimate with small $L^2$ error (averaged across timesteps), we provide efficient convergence guarantees for any data distribution with second-order…
The Kullback-Leibler (KL) divergence plays a central role in probabilistic machine learning, where it commonly serves as the canonical loss function. Optimization in such settings is often performed over the probability simplex, where the…
We prove a general theorem to bound the total variation distance between the distribution of an integer valued random variable of interest and an appropriate discretized normal distribution. We apply the theorem to 2-runs in a sequence of…
The K-sample testing problem involves determining whether K groups of data points are each drawn from the same distribution. Analysis of variance is arguably the most classical method to test mean differences, along with several recent…
Variational Inference approximates an unnormalized distribution via the minimization of Kullback-Leibler (KL) divergence. Although this divergence is efficient for computation and has been widely used in applications, it suffers from some…
Let $(X_{i}, i\in J)$ be a family of locally dependent nonnegative integer-valued random variables, and consider the sum $W=\sum\nolimits_{i\in J}X_i$. We first establish a general error upper bound for $d_{TV}(W, M)$ using Stein's method,…
We derive a Gaussian Central Limit Theorem for the sample quantiles based on locally dependent random variables with explicit convergence rate. Our approach is based on converting the problem to a sum of indicator random variables, applying…