Related papers: Limit Distribution for Smooth Total Variation and …
We develop a unified mathematical framework for certified Top-$k$ attention truncation that quantifies approximation error at both the distribution and output levels. For a single attention distribution $P$ and its Top-$k$ truncation $\hat…
Many decision problems in science, engineering and economics are affected by uncertain parameters whose distribution is only indirectly observable through samples. The goal of data-driven decision-making is to learn a decision from finitely…
We consider the problem of estimating Ising models over $n$ variables in Total Variation (TV) distance, given $l$ independent samples from the model. While the statistical complexity of the problem is well-understood [DMR20], identifying…
Comparing probability distributions is at the crux of many machine learning algorithms. Maximum Mean Discrepancies (MMD) and Wasserstein distances are two classes of distances between probability distributions that have attracted abundant…
Stochastic programs where the uncertainty distribution must be inferred from noisy data samples are considered. The stochastic programs are approximated with distributionally-robust optimizations that minimize the worst-case expected cost…
Score-based diffusion models have demonstrated remarkable empirical success in learning high-dimensional distributions, particularly those exhibiting low-dimensional and multi-modal structures. However, theoretical understanding of their…
This paper introduces a novel approach to securing machine learning model deployments against potential distribution shifts in practical applications, the Total Variation Out-of-Distribution (TV-OOD) detection method. Existing methods have…
Despite the remarkable empirical success of score-based diffusion models, their statistical guarantees remain underdeveloped. Existing analyses often provide pessimistic convergence rates that do not reflect the intrinsic low-dimensional…
Training machine learning and statistical models often involves optimizing a data-driven risk criterion. The risk is usually computed with respect to the empirical data distribution, but this may result in poor and unstable out-of-sample…
A fundamental problem in statistics is estimating the shape matrix of an Elliptical distribution. This generalizes the familiar problem of Gaussian covariance estimation, for which the sample covariance achieves optimal estimation error.…
Mixture models, such as Gaussian mixture models, are widely used in machine learning to represent complex data distributions. A key challenge, especially in high-dimensional settings, is to determine the mixture order and estimate the…
Consider a random sample of $n$ independently and identically distributed $p$-dimensional normal random vectors. A test statistic for complete independence of high-dimensional normal distributions, proposed by Schott (2005), is defined as…
Gaussian processes are distributions over functions that are versatile and mathematically convenient priors in Bayesian modelling. However, their use is often impeded for data with large numbers of observations, $N$, due to the cubic (in…
We develop a projected Wasserstein distance for the two-sample test, a fundamental problem in statistics and machine learning: given two sets of samples, to determine whether they are from the same distribution. In particular, we aim to…
We give necessary and sufficient conditions to characterize the convergence in distribution of a sequence of arbitrary random variables to a probability distribution which is the invariant measure of a diffusion process. This class of…
In this paper, we aim to study the diffusion approximation for multi-scale McKean-Vlasov stochastic differential equations. More precisely, we prove the weak convergence of slow process $X^\varepsilon$ in $C([0,T];\mathbb{R}^n)$ towards the…
This article studies the infinite-width limit of deep feedforward neural networks whose weights are dependent, and modelled via a mixture of Gaussian distributions. Each hidden node of the network is assigned a nonnegative random variable…
We establish a general inequality on the Poisson space, yielding an upper bound for the distance in total variation between the law of a regular random variable with values in the integers and a Poisson distribution. Several applications…
For an ergodic Brownian diffusion with invariant measure $\nu$, we consider a sequence of empirical distributions ($\nu$n) n$\ge$1 associated with an approximation scheme with decreasing time step ($\gamma$n) n$\ge$1 along an adapted…
Suppose $n$ independent random variables $X_1, X_2, \dots, X_n$ have zero mean and equal variance. We prove that if the average of $\chi^2$ distances between these variables and the normal distribution is bounded by a sufficiently small…