Related papers: Edgeworth correction for the largest eigenvalue in…
Consider large signal-plus-noise data matrices of the form $S + \Sigma^{1/2} X$, where $S$ is a low-rank deterministic signal matrix and the noise covariance matrix $\Sigma$ can be anisotropic. We establish the asymptotic joint distribution…
Principal Component Analysis (PCA) is an efficient tool to optimize the multiparameter tests of general relativity (GR) where one tests for simultaneous deviations in multiple post-Newtonian (PN) phasing coefficients by introducing…
We study the classification problem for high-dimensional data with $n$ observations on $p$ features where the $p \times p$ covariance matrix $\Sigma$ exhibits a spiked eigenvalue structure and the vector $\zeta$, given by the difference…
We study the eigenvector mass distribution of an $N\times N$ Wigner matrix on a set of coordinates $I$ satisfying $| I | \ge c N$ for some constant $c >0$. For eigenvectors corresponding to eigenvalues at the spectral edge, we show that the…
We consider large complex random sample covariance matrices obtained from "spiked populations", that is when the true covariance matrix is diagonal with all but finitely many eigenvalues equal to one. We investigate the limiting behavior of…
Non-linear gravitational collapse introduces non-Gaussian statistics into the matter fields of the late Universe. As the large-scale structure is the target of current and future observational campaigns, one would ideally like to have the…
The article considers an inhomogeneous Erd\H{o}s-R\"enyi random graph on $\{1,\ldots, N\}$, where an edge is placed between vertices $i$ and $j$ with probability $\varepsilon_N f(i/N,j/N)$, for $i\le j$, the choice being made independent…
Estimating the leading principal components of data, assuming they are sparse, is a central task in modern high-dimensional statistics. Many algorithms were developed for this sparse PCA problem, from simple diagonal thresholding to…
We consider the problem of regression learning for deterministic design and independent random errors. We start by proving a sharp PAC-Bayesian type bound for the exponentially weighted aggregate (EWA) under the expected squared empirical…
We develop asymptotic theory for principal component analysis (PCA) of a high-dimensional factor model in which the working dimension $R$ is fixed and only required to satisfy $R \ge r$, where $r$ is the true number of factors. Building on…
When training neural networks with full-batch gradient descent (GD) and step size $\eta$, the largest eigenvalue of the Hessian -- the sharpness $S(\boldsymbol{\theta})$ -- rises to $2/\eta$ and hovers there, a phenomenon termed the Edge of…
The expectation-maximization (EM) algorithm is an iterative method for finding maximum likelihood estimates when data are incomplete or are treated as being incomplete. The EM algorithm and its variants are commonly used for parameter…
Edgeworth expansion provides higher-order corrections to the normal approximation for a probability distribution. The classical proof of Edgeworth expansion is via characteristic functions. As a powerful method for distributional…
We establish a large-deviations principle for the largest eigenvalue of a generalized sample covariance matrix, meaning a matrix proportional to $Z^T \Gamma Z$, where $Z$ has i.i.d. real or complex entries and $\Gamma$ is not necessarily…
In this paper, we characterize the asymptotic and large scale behavior of the eigenvalues of wavelet random matrices in high dimensions. We assume that possibly non-Gaussian, finite-variance $p$-variate measurements are made of a…
In this paper, we consider the sphericity test for a one-sample problem under high-dimensional two-step monotone incomplete data. Existing asymptotic expansions for the null distributions of the likelihood ratio test (LRT) statistic and…
We consider sample covariance matrices of the form $\mathcal{Q}=(\Sigma^{1/2}X)(\Sigma^{1/2} X)^*$, where the sample $X$ is an $M\times N$ random matrix whose entries are real independent random variables with variance $1/N$ and where…
We study general singular value shrinkage estimators in high-dimensional regression and classification, when the number of features and the sample size both grow proportionally to infinity. We allow models with general covariance matrices…
The spiked covariance model has gained increasing popularity in high-dimensional data analysis. A fundamental problem is determination of the number of spiked eigenvalues, $K$. For estimation of $K$, most attention has focused on the use of…
A large i.i.d. random matrix with deterministic low-rank perturbation has been extensively studied, particularly in the aspects of the ESD (Empirical Spectral Distribution) and the outliers of eigenvalues. In this work, we investigate the…