Related papers: Distortion mismatch in the quantization of probabi…
Many authors have studied the phenomenon of typically Gaussian marginals of high-dimensional random vectors; e.g., for a probability measure on $\R^d$, under mild conditions, most one-dimensional marginals are approximately Gaussian if $d$…
We consider the problem of rate/distortion with side information available only at the decoder. For the case of jointly-Gaussian source X and side information Y, and mean-squared error distortion, Wyner proved in 1976 that the…
This paper studies the probabilistic function approximation problem over reproducing kernel Hilbert spaces. We show the existence and uniqueness of the optimizer under mild assumptions. Furthermore, we generalize the celebrated representer…
We study the problem of distributed mean estimation and optimization under communication constraints. We propose a correlated quantization protocol whose leading term in the error guarantee depends on the mean deviation of data points…
We study the least squares estimator in the residual variance estimation context. We show that the mean squared differences of paired observations are asymptotically normally distributed. We further establish that, by regressing the mean…
In this work, we extend the classical framework of quantization for Borel probability measures defined on normed spaces $\mathbb{R}^k$ by introducing and analyzing the notions of the $n$th constrained quantization error, constrained…
Datasets from the fields of bioinformatics, chemometrics, and face recognition are typically characterized by small samples of high-dimensional data. Among the many variants of linear discriminant analysis that have been proposed in order…
We propose and analyse randomized cubature formulae for the numerical integration of functions with respect to a given probability measure $\mu$ defined on a domain $\Gamma \subseteq \mathbb{R}^d$, in any dimension $d$. Each cubature…
We present novel bounds for estimating discrete probability distributions under the $\ell_\infty$ norm. These are nearly optimal in various precise senses, including a kind of instance-optimality. Our data-dependent convergence guarantees…
Let $E$ be a Moran set on $\mathbb{R}^1$ associated with a closed interval $J$ and two sequences $(n_k)_{k=1}^\infty$ and $(\mathcal{C}_k=(c_{k,j})_{j=1}^{n_k})_{k\geq1}$. Let $\mu$ be the infinite product measure (Moran measure) on $E$…
Let $\{f_i\}_{i=1}^N$ be a set of equi-contractive similitudes on $\mathbb{R}^1$ satisfying the finite-type condition. We study the asymptotic quantization error for self-similar measures $\mu$ associated with $\{f_i\}_{i=1}^N$ and a…
Several researchers have proposed minimisation of maximum mean discrepancy (MMD) as a method to quantise probability measures, i.e., to approximate a target distribution by a representative point set. We consider sequential algorithms that…
We derive limiting distributions of symmetrized estimators of scatter, where instead of all $n(n-1)/2$ pairs of the $n$ observations we only consider $nd$ suitably chosen pairs, $1 \le d < \lfloor n/2\rfloor$. It turns out that the…
Modern machine learning embeddings provide powerful compression of high-dimensional data, yet they typically destroy the geometric structure required for classical likelihood-based statistical inference. This paper develops a rigorous…
We show that the maximum expected inner product between a random vector and the standard normal vector over all couplings subject to a mutual information constraint or regularization is equivalent to a truncated integral involving the…
Let $P$ be a Borel probability measure on $\mathbb R^2$ supported by the Cantor dusts generated by a set of $4^u,\ u\geq 1$, contractive similarity mappings satisfying the strong separation condition. For this probability measure, we…
Let $X_1,\dots,X_n$ be i.i.d. log-concave random vectors in $\mathbb R^d$ with mean 0 and covariance matrix $\Sigma$. We study the problem of quantifying the normal approximation error for $W=n^{-1/2}\sum_{i=1}^nX_i$ with explicit…
The problem of estimating the Kullback-Leibler divergence $D(P\|Q)$ between two unknown distributions $P$ and $Q$ is studied, under the assumption that the alphabet size $k$ of the distributions can scale to infinity. The estimation is…
We consider a linear ill-posed equation in the Hilbert space setting. Multiple independent unbiased measurements of the right hand side are available. A natural approach is to take the average of the measurements as an approximation of the…
This expository paper provides a unified and pedagogical introduction to optimal quantization for probability measures supported on spherical curves and discrete subsets of the sphere, emphasizing both continuous and discrete settings. We…