Related papers: Nearly optimal central limit theorem and bootstrap…
Let C = C(l_1, ..., l_n) be the n-dimensional orthogonal cross-polytope whose axes are of length l_1,..., l_n. Subject to the condition \sum l_i^2 = 1, the mean width of C is minimised when l_i = 1/sqrt{n} for every i, and it is maximised…
We use a new method via $p$-Wasserstein bounds to prove Cram\'er-type moderate deviations in (multivariate) normal approximations. In the classical setting that $W$ is a standardized sum of $n$ independent and identically distributed…
Let $d\geq 3$ be fixed and $G$ be a large random $d$-regular graph on $n$ vertices. We show that if $n$ is large enough then the entry distribution of every almost eigenvector $v$ of $G$ (with entry sum 0 and normalized to have length…
The aim of this note is to investigate the Kolmogorov distance of the Circular Law to the empirical spectral distribution of non-Hermitian random matrices with independent entries. The optimal rate of convergence is determined by the…
We study the minimal sample size N=N(n) that suffices to estimate the covariance matrix of an n-dimensional distribution by the sample covariance matrix in the operator norm, with an arbitrary fixed accuracy. We establish the optimal bound…
We study the fundamental problem of high-dimensional mean estimation in a robust model where a constant fraction of the samples are adversarially corrupted. Recent work gave the first polynomial time algorithms for this problem with…
We consider the problem of optimal approximation of a target measure by an atomic measure with $N$ atoms, in branched optimal transport distance. This is a new branched transport version of optimal quantization problems. New difficulties…
We revisit the null distribution of the high-dimensional spatial-sign test of Wang et al. (2015) under mild structural assumptions on the scatter matrix. We show that the standardized test statistic converges to a non-Gaussian limit,…
The random intersection graph model $\mathcal G(n,m,p)$ is considered. Due to substantial edge dependencies, studying even fundamental statistics such as the subgraph count is significantly more challenging than in the classical binomial…
We study the problem of high-dimensional robust linear regression where a learner is given access to $n$ samples from the generative model $Y = \langle X,w^* \rangle + \epsilon$ (with $X \in \mathbb{R}^d$ and $\epsilon$ independent), in…
We prove that the distribution density of any non-constant polynomial $f(\xi_1,\xi_2,\ldots)$ of degree $d$ in independent standard Gaussian random variables $\xi$ (possibly, in infinitely many variables) always belongs to the…
Given a vector $F=(F_1,\dots,F_m)$ of Poisson functionals $F_1,\dots,F_m$, we investigate the proximity between $F$ and an $m$-dimensional centered Gaussian random vector $N_\Sigma$ with covariance matrix $\Sigma\in\mathbb{R}^{m\times m}$.…
Differential entropy and log determinant of the covariance matrix of a multivariate Gaussian distribution have many applications in coding, communications, signal processing and statistical inference. In this paper we consider in the high…
Bootstrap smoothed (bagged) parameter estimators have been proposed as an improvement on estimators found after preliminary data-based model selection. The key result of Efron (2014) is a very convenient and widely applicable formula for a…
This paper studies the problem of estimating a covariance matrix from correlated sub-Gaussian samples. We consider using the correlated sample covariance matrix estimator to approximate the true covariance matrix. We establish…
Linear thresholding models postulate that the conditional distribution of a response variable in terms of covariates differs on the two sides of a (typically unknown) hyperplane in the covariate space. A key goal in such models is to learn…
Using entropic inequalities from information theory, we provide new bounds on the total variation and 2-Wasserstein distances between a conditionally Gaussian law and a Gaussian law with invertible covariance matrix. We apply our results to…
We consider distributed estimation of the inverse covariance matrix, also called the concentration or precision matrix, in Gaussian graphical models. Traditional centralized estimation often requires global inference of the covariance…
This work develops algorithms for non-parametric confidence regions for samples from a univariate distribution whose support is a discrete mesh bounded on the left. We generalize the theory of Learned-Miller to preorders over the sample…
A limit theorem for the largest interpoint distance of $p$ independent and identically distributed points in $\mathbb{R}^n$ to the Gumbel distribution is proved, where the number of points $p=p_n$ tends to infinity as the dimension of the…