Related papers: On the tensorization of the variational distance
With the proliferation of generative AI and the increasing volume of generative data (also called as synthetic data), assessing the fidelity of generative data has become a critical concern. In this paper, we propose a discriminative…
Statistical divergences are ubiquitous in machine learning as tools for measuring discrepancy between probability distributions. As these applications inherently rely on approximating distributions from samples, we consider empirical…
We use a new method via $p$-Wasserstein bounds to prove Cram\'er-type moderate deviations in (multivariate) normal approximations. In the classical setting that $W$ is a standardized sum of $n$ independent and identically distributed…
Given i.i.d.~samples from an unknown distribution $P$, the goal of distribution learning is to recover the parameters of a distribution that is close to $P$. When $P$ belongs to the class of product distributions on the Boolean hypercube…
We develop a general method for lower bounding the variance of sequences in arithmetic progressions mod $q$, summed over all $q \leq Q$, building on previous work of Liu, Perelli, Hooley, and others. The proofs lower bound the variance by…
We develop a unified mathematical framework for certified Top-$k$ attention truncation that quantifies approximation error at both the distribution and output levels. For a single attention distribution $P$ and its Top-$k$ truncation $\hat…
We provide new upper and lower bounds on the minimum possible ratio of the spectral and Frobenius norms of a (partially) symmetric tensor. In the particular case of general tensors our result recovers a known upper bound. For symmetric…
For each $n\ge 1$, let $X_{n,1},\ldots,X_{n,N_n}$ be real random variables and $S_n=\sum_{i=1}^{N_n}X_{n,i}$. Let $m_n\ge 1$ be an integer. Suppose $(X_{n,1},\ldots,X_{n,N_n})$ is $m_n$-dependent, $E(X_{ni})=0$, $E(X_{ni}^2)<\infty$ and…
We develop a new formulation of Stein's method to obtain computable upper bounds on the total variation distance between the geometric distribution and a distribution of interest. Our framework reduces the problem to the construction of a…
The first part of this work considers the entropy of the sum of (possibly dependent and non-identically distributed) Bernoulli random variables. Upper bounds on the error that follows from an approximation of this entropy by the entropy of…
In this paper, we consider using total variation minimization to recover signals whose gradients have a sparse support, from a small number of measurements. We establish the proof for the performance guarantee of total variation (TV)…
Pinsker's classical inequality asserts that the total variation $TV(\mu, \nu)$ between two probability measures is bounded by $\sqrt{ 2H(\mu|\nu)}$ where $H$ denotes the relative entropy (or Kullback-Leibler divergence). Considering the…
Uniform continuity bounds on entropies are generally expressed in terms of a single distance measure between a pair of probability distributions or quantum states, typically, the total variation distance or trace distance. However, if an…
A pair of probability distributions over $\{0,1\}^n$ is said to be $(k,\delta)$-wise indistinguishable if all of the size $k$ marginals are within statistical distance at most $\delta$. Previous works introduced this concept and study when…
We give a simple polynomial-time approximation algorithm for the total variation distance between two product distributions.
In this paper, we establish a novel connection between total variation (TV) distance estimation and probabilistic inference. In particular, we present an efficient, structure-preserving reduction from relative approximation of TV distance…
This work considers the use of Total variation (TV) minimization in the recovery of a given gradient sparse vector from Gaussian linear measurements. It has been shown in recent studies that there exist a sharp phase transition behavior in…
Upper bounds on the Kolmogorov distance (and, equivalently in this case, on the total variation distance) between the Student distribution with p degrees of freedom (SD_p) and the standard normal distribution are obtained. These bounds are…
We prove a variational principle for the upper and lower metric mean dimension of level sets \[ \left\{x\in X: \lim_{n\to\infty}\frac{1}{n}\sum_{j=0}^{n-1}\varphi(f^{j}(x))=\alpha\right\} \] associated to continuous potentials $\varphi:X\to…
Total variation (TV) minimization is one of the most important techniques in modern signal/image processing, and has wide range of applications. While there are numerous recent works on the restoration guarantee of the TV minimization in…