Related papers: Edgeworth correction for the largest eigenvalue in…
Modern datasets are trending towards ever higher dimension. In response, recent theoretical studies of covariance estimation often assume the proportional-growth asymptotic framework, where the sample size $n$ and dimension $p$ are…
In this article, we establish a limiting distribution for eigenvalues of a class of auto-covariance matrices. The same distribution has been found in the literature for a regularized version of these auto-covariance matrices. The original…
The neighbourhood of the largest eigenvalue $\lambda_{\rm max}$ in the Gaussian unitary ensemble (GUE) and Laguerre unitary ensemble (LUE) is referred to as the soft edge. It is known that there exists a particular centring and scaling such…
Given a matrix of observed data, Principal Components Analysis (PCA) computes a small number of orthogonal directions that contain most of its variability. Provably accurate solutions for PCA have been in use for a long time. However, to…
Given a matrix of observed data, Principal Components Analysis (PCA) computes a small number of orthogonal directions that contain most of its variability. Provably accurate solutions for PCA have been in use for a long time. However, to…
In this paper we focus on the finite n probability distribution function of the largest eigenvalue in the classical Gaussian Ensemble of n by n matrices (GEn). We derive the finite n largest eigenvalue probability distribution function for…
We consider a measure $\psi$ k of dispersion which extends the notion of Wilk's generalised variance, or entropy, for a d-dimensional distribution, and is based on the mean squared volume of simplices of dimension k $\le$ d formed by k + 1…
Principal Component Analysis (PCA) is a cornerstone of dimensionality reduction, yet its classical formulation relies critically on second-order moments and is therefore fragile in the presence of heavy-tailed data and impulsive noise.…
Low-rank matrix recovery problems involving high-dimensional and heterogeneous data appear in applications throughout statistics and machine learning. The contribution of this paper is to establish the fundamental limits of recovery for a…
The top eigenvalues of rank $r$ spiked real Wishart matrices and additively perturbed Gaussian orthogonal ensembles are known to exhibit a phase transition in the large size limit. We show that they have limiting distributions for…
In the context of principal components analysis (PCA), the bootstrap is commonly applied to solve a variety of inference problems, such as constructing confidence intervals for the eigenvalues of the population covariance matrix $\Sigma$.…
Sample covariance matrices from multi-population typically exhibit several large spiked eigenvalues, which stem from differences between population means and are crucial for inference on the underlying data structure. This paper…
We present a novel technique for sparse principal component analysis. This method, named Eigenvectors from Eigenvalues Sparse Principal Component Analysis (EESPCA), is based on the formula for computing squared eigenvector loadings of a…
In this paper, we show that the largest and smallest eigenvalues of a sample correlation matrix stemming from $n$ independent observations of a $p$-dimensional time series with iid components converge almost surely to $(1+\sqrt{\gamma})^2$…
We establish the validity of the empirical Edgeworth expansion (EE) for a studentized trimmed mean, under the sole condition that the underlying distribution function of the observations satisfies a local smoothness condition near the two…
Previous versions of sparse principal component analysis (PCA) have presumed that the eigen-basis (a $p \times k$ matrix) is approximately sparse. We propose a method that presumes the $p \times k$ matrix becomes approximately sparse after…
How do statistical dependencies in measurement noise influence high-dimensional inference? To answer this, we study the paradigmatic spiked matrix model of principal components analysis (PCA), where a rank-one matrix is corrupted by…
We develop generalized approach to obtaining Edgeworth expansions for $t$-statistics of an arbitrary order using computer algebra and combinatorial algorithms. To incorporate various versions of mean-based statistics, we introduce Adjusted…
Probabilistic principal component analysis (PCA) and its Bayesian variant (BPCA) are widely used for dimension reduction in machine learning and statistics. The main advantage of probabilistic PCA over the traditional formulation is…
We investigate the statistics of the largest eigenvalue, $\lambda_{\rm max}$, in an ensemble of $N\times N$ large ($N\gg 1$) sparse adjacency matrices, $A_N$. The most attention is paid to the distribution and typical fluctuations of…