Related papers: Sampling without replacement from a high-dimension…
Given a large, high-dimensional sample from a spiked population, the top sample covariance eigenvalue is known to exhibit a phase transition. We show that the largest eigenvalues have asymptotic distributions near the phase transition in…
The Davis--Kahan theorem is used in the analysis of many statistical procedures to bound the distance between subspaces spanned by population eigenvectors and their sample versions. It relies on an eigenvalue separation condition between…
We consider testing for two-sample means of high dimensional populations by thresholding. Two tests are investigated, which are designed for better power performance when the two population mean vectors differ only in sparsely populated…
We propose a new randomized optimization method for high-dimensional problems which can be seen as a generalization of coordinate descent to random subspaces. We show that an adaptive sampling strategy for the random subspace significantly…
We provide some asymptotic theory for the largest eigenvalues of a sample covariance matrix of a p-dimensional time series where the dimension p = p_n converges to infinity when the sample size n increases. We give a short overview of the…
This paper considers the problem of estimating the population spectral distribution from a sample covariance matrix in large dimensional situations. We generalize the contour-integral based method in Mestre (2008) and present a local moment…
We show that the fluctuations of the largest eigenvalue of any generalized Wigner matrix $H$ converge to the Tracy-Widom laws at a rate nearly $O(N^{-1/3})$, as the matrix dimension $N$ tends to infinity. We allow the variances of the…
We review old and recent finite de Finetti theorems in total variation distance and in relative entropy, and we highlight their connections with bounds on the difference between sampling with and without replacement. We also establish two…
A methodology to analyze the properties of the first (largest) eigenvalue and its eigenvector is developed for large symmetric random sparse matrices utilizing the cavity method of statistical mechanics. Under a tree approximation, which is…
Let X be a n*p matrix and l_1 the largest eigenvalue of the covariance matrix X^{*}*X. The "null case" where X_{i,j} are independent Normal(0,1) is of particular interest for principal component analysis. For this model, when n, p tend to…
We propose novel methodology for testing equality of model parameters between two high-dimensional populations. The technique is very general and applicable to a wide range of models. The method is based on sample splitting: the data is…
In this paper, we derive the explicit series expansion of the eigenvalue distribution of various models, namely the case of non-central Wishart distributions, as well as correlated zero mean Wishart distributions. The tools used extend…
We develop an estimator for the high-dimensional covariance matrix of a locally stationary process with a smoothly varying trend and use this statistic to derive consistent predictors in non-stationary time series. In contrast to the…
Given two positive integers $n$ and $k$ and a parameter $t\in (0,1)$, we choose at random a vector subspace $V_{n}\subset \mathbb{C}^{k}\otimes\mathbb{C}^{n}$ of dimension $N\sim tnk$. We show that the set of $k$-tuples of singular values…
We investigate the asymptotics of eigenvalues of sample covariance matrices associated with a class of non-independent Gaussian processes (separable and temporally stationary) under the Kolmogorov asymptotic regime. The limiting spectral…
Sample correlation matrices are employed ubiquitously in statistics. However, quite surprisingly, little is known about their asymptotic spectral properties for high-dimensional data, particularly beyond the case of "null models" for which…
The correlated Wishart model provides a standard tool for the analysis of correlations in a rich variety of systems. Although much is known for complex correlation matrices, the empirically much more important real case still poses…
In recent years, sparse principal component analysis has emerged as an extremely popular dimension reduction technique for high-dimensional data. The theoretical challenge, in the simplest case, is to estimate the leading eigenvector of a…
Estimating the expected value of an observable appearing in a non-equilibrium stochastic process usually involves sampling. If the observable's variance is high, many samples are required. In contrast, we show that performing the same task…
We study the space spanned by the integer shifts of a bivariate Gaussian function and the problem of reconstructing any function in that space from samples scattered across the plane. We identify a large class of lattices, or more generally…