Related papers: Edgeworth correction for the largest eigenvalue in…
We show that in a common high-dimensional covariance model, the choice of loss function has a profound effect on optimal estimation. In an asymptotic framework based on the Spiked Covariance model and use of orthogonally invariant…
In this paper, the key objects of interest are the sequential covariance matrices $\mathbf{S}_{n,t}$ and their largest eigenvalues. Here, the matrix $\mathbf{S}_{n,t}$ is computed as the empirical covariance associated with observations…
A large class of statistics can be formulated as smooth functions of sample means of random vectors. In this paper, we propose a general partial Cram\'{e}r's condition (GPCC) and apply it to establish the validity of the Edgeworth expansion…
We analyze the prediction error of principal component regression (PCR) and prove high probability bounds for the corresponding squared risk conditional on the design. Our first main result shows that PCR performs comparably to the oracle…
Principal Component Analysis (PCA) is a widely used method for dimensionality reduction, but it often overlooks fairness, especially when working with data that includes demographic characteristics. This can lead to biased representations…
We study principal components analyses in multivariate random and mixed effects linear models, assuming a spherical-plus-spikes structure for the covariance matrix of each random effect. We characterize the behavior of outlier sample…
This work studies estimation of sparse principal components in high dimensions. Specifically, we consider a class of estimators based on kernel PCA, generalizing the covariance thresholding algorithm proposed by Krauthgamer et al. (2015).…
Sparse principal component analysis (sparse PCA) is a widely used technique for dimensionality reduction in multivariate analysis, addressing two key limitations of standard PCA. First, sparse PCA can be implemented in high-dimensional low…
Principal component analysis (PCA) is a key tool in the field of data dimensionality reduction that is useful for various data science problems. However, many applications involve heterogeneous data that varies in quality due to noise…
In this article we generalize the classical Edgeworth expansion for the probability density function (PDF) of sums of a finite number of symmetric independent identically distributed random variables with a finite variance to sums of…
We show how to calculate individual terms of the Edgeworth series to approximate the distribution of the Pearson correlation coefficient with the help of a simple Mathematica program. We also demonstrate how to eliminate the corresponding…
We consider matrices formed by a random $N\times N$ matrix drawn from the Gaussian Orthogonal Ensemble (or Gaussian Unitary Ensemble) plus a rank-one perturbation of strength $\theta$, and focus on the largest eigenvalue, $x$, and the…
We present a method for performing Principal Component Analysis (PCA) on noisy datasets with missing values. Estimates of the measurement error are used to weight the input data such that compared to classic PCA, the resulting eigenvectors…
We generalize the maximum likelihood method to non-Gaussian distribution functions by means of the multivariate Edgeworth expansion. We stress the potential interest of this technique in all those cosmological problems in which the…
Principal Component Analysis (PCA) is one of the most commonly used statistical methods for data exploration, and for dimensionality reduction wherein the first few principal components account for an appreciable proportion of the…
We consider the limiting location and limiting distribution of the largest eigenvalue in real symmetric ($\beta$ = 1), Hermitian ($\beta$ = 2), and Hermitian self-dual ($\beta$ = 4) random matrix models with rank 1 external source. They are…
In statistics and machine learning, people are often interested in the eigenvectors (or singular vectors) of certain matrices (e.g. covariance matrices, data matrices, etc). However, those matrices are usually perturbed by noises or…
We consider principal component analysis (PCA) in decomposable Gaussian graphical models. We exploit the prior information in these models in order to distribute its computation. For this purpose, we reformulate the problem in the sparse…
In this paper, we consider a data matrix $X_N\in\mathbb{R}^{N\times p}$ where all the rows are i.i.d. samples in $\mathbb{R}^p$ of mean zero and covariance matrix $\Sigma\in\mathbb{R}^{p\times p}$. Here the population matrix $\Sigma$ is of…
Principal Component Analysis (PCA) is a dimension reduction technique. It produces inconsistent estimators when the dimensionality is moderate to high, which is often the problem in modern large-scale applications where algorithm…