English

Classification of high-dimensional data with spiked covariance matrix structure

Machine Learning 2026-02-12 v3 Machine Learning

Abstract

We study the classification problem for high-dimensional data with nn observations on pp features where the p×pp \times p covariance matrix Σ\Sigma exhibits a spiked eigenvalue structure and the vector ζ\zeta, given by the difference between the {\em whitened} mean vectors, is sparse. We analyze an adaptive classifier (adaptive with respect to the sparsity ss) that first performs dimension reduction on the feature vectors prior to classification in the dimensionally reduced space, i.e., the classifier whitens the data, then screens the features by keeping only those corresponding to the ss largest coordinates of ζ\zeta and finally applies Fisher linear discriminant on the selected features. Leveraging recent results on entrywise matrix perturbation bounds for covariance matrices, we show that the resulting classifier is Bayes optimal whenever nn \rightarrow \infty and sn1lnp0s \sqrt{n^{-1} \ln p} \rightarrow 0. Notably, our theory also guarantees Bayes optimality for the corresponding quadratic discriminant analysis (QDA). Experimental results on real and synthetic data further indicate that the proposed approach is competitive with state-of-the-art methods while operating on a substantially lower-dimensional representation.

Keywords

Cite

@article{arxiv.2110.01950,
  title  = {Classification of high-dimensional data with spiked covariance matrix structure},
  author = {Yin-Jen Chen and Minh Tang},
  journal= {arXiv preprint arXiv:2110.01950},
  year   = {2026}
}

Comments

40 pages, 2 figures

R2 v1 2026-06-24T06:37:53.544Z