English
Related papers

Related papers: Classification of high-dimensional data with spike…

200 papers

While the SLIM approach obtained high ranking-accuracy in many experiments in the literature, it is also known for its high computational cost of learning its parameters from data. For this reason, we focus in this paper on variants of…

Information Retrieval · Computer Science 2019-05-01 Harald Steck

We show that in a common high-dimensional covariance model, the choice of loss function has a profound effect on optimal estimation. In an asymptotic framework based on the Spiked Covariance model and use of orthogonally invariant…

Statistics Theory · Mathematics 2017-06-06 David L. Donoho , Matan Gavish , Iain M. Johnstone

In the field of statistical learning and data analysis, estimating precision matrices (i.e., the inverse of covariance matrices) is a critical task, particularly for understanding dependency structures among variables. However, traditional…

Methodology · Statistics 2026-05-15 Zhongfeng Qin , Hao Xu , Wenhao Cui , Wan Tian

Linear discriminant analysis (LDA) is a classical method for dimensionality reduction, where discriminant vectors are sought to project data to a lower dimensional space for optimal separability of classes. Several recent papers have…

Computation · Statistics 2022-03-04 Summer Atkins , Gudmundur Einarsson , Brendan Ames , Line Clemmensen

This paper investigates the fundamental limits for detecting a high-dimensional sparse matrix contaminated by white Gaussian noise from both the statistical and computational perspectives. We consider $p\times p$ matrices whose rows and…

Statistics Theory · Mathematics 2018-01-03 T. Tony Cai , Yihong Wu

This paper addresses the problem of inverse covariance (also known as precision matrix) estimation in high-dimensional settings. Specifically, we focus on two classes of estimators: linear shrinkage estimators with a target proportional to…

Machine Learning · Statistics 2025-11-21 Lucas Morisset , Adrien Hardy , Alain Durmus

In this paper, the key objects of interest are the sequential covariance matrices $\mathbf{S}_{n,t}$ and their largest eigenvalues. Here, the matrix $\mathbf{S}_{n,t}$ is computed as the empirical covariance associated with observations…

Statistics Theory · Mathematics 2024-05-01 Nina Dörnemann , Debashis Paul

We introduce a new algorithm, called adaptive sparse backfitting algorithm, for solving high dimensional Sparse Additive Model (SpAM) utilizing symmetric, non-negative definite smoothers. Unlike the previous sparse backfitting algorithm,…

Machine Learning · Statistics 2014-11-13 Yan Li

A generalized spiked Fisher matrix is considered in this paper. We establish a criterion for the description of the support of the limiting spectral distribution of high-dimensional generalized Fisher matrix and study the almost sure limits…

Statistics Theory · Mathematics 2019-12-09 Dandan Jiang , Jiang Hu , Zhiqiang Hou

We consider the task of classification in the high dimensional setting where the number of features of the given data is significantly greater than the number of observations. To accomplish this task, we propose a heuristic, called sparse…

Machine Learning · Statistics 2015-12-09 Brendan P. W. Ames , Mingyi Hong

We address a frequently asked question on the covariance fitting of the highly correlated data such as our $B_K$ data based on the SU(2) staggered chiral perturbation theory. Basically, the essence of the problem is that we do not have an…

High Energy Physics - Lattice · Physics 2012-08-22 Boram Yoon , Yong-Chull Jang , Chulwoo Jung , Weonjong Lee

Feature selection (FS) has become an indispensable task in dealing with today's highly complex pattern recognition problems with massive number of features. In this study, we propose a new wrapper approach for FS based on binary…

Machine Learning · Statistics 2016-03-08 Vural Aksakalli , Milad Malekipirbazari

How do statistical dependencies in measurement noise influence high-dimensional inference? To answer this, we study the paradigmatic spiked matrix model of principal components analysis (PCA), where a rank-one matrix is corrupted by…

Information Theory · Computer Science 2023-06-05 Jean Barbier , Francesco Camilli , Marco Mondelli , Manuel Saenz

We develop a method for estimating well-conditioned and sparse covariance and inverse covariance matrices from a sample of vectors drawn from a sub-gaussian distribution in high dimensional setting. The proposed estimators are obtained by…

Statistics Theory · Mathematics 2016-11-21 Ashwini Maurya

This paper considers statistical inference for the explained variance $\beta^{\intercal}\Sigma \beta$ under the high-dimensional linear model $Y=X\beta+\epsilon$ in the semi-supervised setting, where $\beta$ is the regression vector and…

Methodology · Statistics 2020-12-01 T. Tony Cai , Zijian Guo

This paper aims to develop an optimality theory for linear discriminant analysis in the high-dimensional setting. A data-driven and tuning free classification rule, which is based on an adaptive constrained $\ell_1$ minimization approach,…

Methodology · Statistics 2018-04-10 T. Tony Cai , Linjun Zhang

Missing data occur frequently in a wide range of applications. In this paper, we consider estimation of high-dimensional covariance matrices in the presence of missing observations under a general missing completely at random model in the…

Methodology · Statistics 2016-05-17 T. Tony Cai , Anru Zhang

In recent years, sparse principal component analysis has emerged as an extremely popular dimension reduction technique for high-dimensional data. The theoretical challenge, in the simplest case, is to estimate the leading eigenvector of a…

Statistics Theory · Mathematics 2016-09-29 Tengyao Wang , Quentin Berthet , Richard J. Samworth

Topological data analysis (TDA) has emerged as one of the most promising techniques to reconstruct the unknown shapes of high-dimensional spaces from observed data samples. TDA, thus, yields key shape descriptors in the form of persistent…

Machine Learning · Statistics 2017-11-15 Wei Guo , Krithika Manohar , Steven L. Brunton , Ashis G. Banerjee

In this work, we propose a scalable Bayesian procedure for learning the local dependence structure in a high-dimensional model where the variables possess a natural ordering. The ordering of variables can be indexed by time, the vicinities…

Methodology · Statistics 2021-09-27 Kyoungjae Lee , Lizhen Lin