Related papers: A Tracy-Widom Empirical Estimator For Valid P-valu…
Let A be a p-variate real Wishart matrix on n degrees of freedom with identity covariance. The distribution of the largest eigenvalue in A has important applications in multivariate statistics. Consider the asymptotics when p grows in…
The correlated Wishart model provides a standard tool for the analysis of correlations in a rich variety of systems. Although much is known for complex correlation matrices, the empirically much more important real case still poses…
We propose novel methodology for testing equality of model parameters between two high-dimensional populations. The technique is very general and applicable to a wide range of models. The method is based on sample splitting: the data is…
A fundamental problem in multivariate analysis is testing general linear hypotheses for regression coefficients in a multivariate linear model. This framework encompasses a wide range of well-studied tasks, including MANOVA, joint…
Let $A$ and $B$ be independent, central Wishart matrices in $p$ variables with common covariance and having $m$ and $n$ degrees of freedom, respectively. The distribution of the largest eigenvalue of $(A+B)^{-1}B$ has numerous applications…
Under certain conditions, the largest eigenvalue of a sample covariance matrix undergoes a well-known phase transition when the sample size $n$ and data dimension $p$ diverge proportionally. In the subcritical regime, this eigenvalue has…
We propose a method for constructing p-values for general hypotheses in a high-dimensional linear model. The hypotheses can be local for testing a single regression parameter or they may be more global involving several up to all…
In this paper, we develop a systematic theory for high dimensional analysis of variance in multivariate linear regression, where the dimension and the number of coefficients can both grow with the sample size. We propose a new \emph{U}~type…
The greatest root distribution occurs everywhere in classical multivariate analysis, but even under the null hypothesis the exact distribution has required extensive tables or special purpose software. We describe a simple approximation,…
In the present work we have selected a collection of statistical and mathematical tools useful for the exploration of multivariate data and we present them in a form that is meant to be particularly accessible to a classically trained…
Assigning significance in high-dimensional regression is challenging. Most computationally efficient selection algorithms cannot guard against inclusion of noise variables. Asymptotically valid p-values are not available. An exception is a…
We compute the Tracy-Widom distribution describing the asymptotic distribution of the largest eigenvalue of a large random matrix by solving a boundary-value problem posed by Bloemendal in his Ph.D. Thesis (2011). The distribution is…
High-dimensional linear classifiers, such as the support vector machine (SVM) and distance weighted discrimination (DWD), are commonly used in biomedical research to distinguish groups of subjects based on a large number of features.…
The distributions of the largest and the smallest eigenvalues of a $p$-variate sample covariance matrix $S$ are of great importance in statistics. Focusing on the null case where $nS$ follows the standard Wishart distribution $W_p(I,n)$, we…
Let ${\bf X, Y} $ denote two independent real Gaussian $\mathsf{p} \times \mathsf{m}$ and $\mathsf{p} \times \mathsf{n}$ matrices with $\mathsf{m}, \mathsf{n} \geq \mathsf{p}$, each constituted by zero mean i.i.d. columns with common…
In applied multivariate statistics, estimating the number of latent dimensions or the number of clusters, $k$, is a fundamental and recurring problem. We study a sequence of statistics called "cross-validated eigenvalues." Under a large…
Recently Johansson and Johnstone proved that the distribution of the (properly rescaled) largest principal component of the complex (real) Wishart matrix $ X^* \* X (X^t \*X) $ converges to the Tracy-Widom law as $ n, p $ (the dimensions of…
We consider a multivariate linear response regression in which the number of responses and predictors is large and comparable with the number of observations, and the rank of the matrix of regression coefficients is assumed to be small. We…
We derive efficient recursive formulas giving the exact distribution of the largest eigenvalue for finite dimensional real Wishart matrices and for the Gaussian Orthogonal Ensemble (GOE). In comparing the exact distribution with the…
Cross-validation is a statistical tool that can be used to improve large covariance matrix estimation. Although its efficiency is observed in practical applications and a convergence result towards the error of the non linear shrinkage is…