English
Related papers

Related papers: Information-theoretic bounds and phase transitions…

200 papers

We give upper and lower bounds on the information-theoretic threshold for community detection in the stochastic block model. Specifically, let $k$ be the number of groups, $d$ be the average degree, the probability of edges between vertices…

Probability · Mathematics 2016-04-25 Jess Banks , Cristopher Moore

The robust PCA problem, wherein, given an input data matrix that is the superposition of a low-rank matrix and a sparse matrix, we aim to separate out the low-rank and sparse components, is a well-studied problem in machine learning. One…

Machine Learning · Computer Science 2017-07-06 U. N. Niranjan , Arun Rajkumar , Theja Tulabandhula

Factorizing low-rank matrices is a problem with many applications in machine learning and statistics, ranging from sparse PCA to community detection and sub-matrix localization. For probabilistic models in the Bayes optimal setting, general…

Information Theory · Computer Science 2018-12-07 Jean Barbier , Mohamad Dia , Nicolas Macris , Florent Krzakala , Lenka Zdeborová

This paper studies the statistical and computational limits of high-order clustering with planted structures. We focus on two clustering models, constant high-order clustering (CHC) and rank-one higher-order clustering (ROHC), and study the…

Statistics Theory · Mathematics 2023-10-04 Yuetian Luo , Anru R. Zhang

The interplay between computational efficiency and statistical accuracy in high-dimensional inference has drawn increasing attention in the literature. In this paper, we study computational and statistical boundaries for submatrix…

Statistics Theory · Mathematics 2020-07-27 T. Tony Cai , Tengyuan Liang , Alexander Rakhlin

Given full or partial information about a collection of points that lie close to a union of several subspaces, subspace clustering refers to the process of clustering the points according to their subspace and identifying the subspaces. One…

Machine Learning · Statistics 2018-01-16 Zachary Charles , Amin Jalali , Rebecca Willett

We study the statistical decision process of detecting the low-rank signal from various signal-plus-noise type data matrices, known as the spiked random matrix models. We first show that the principal component analysis can be improved by…

Statistics Theory · Mathematics 2023-01-18 Ji Hyung Jung , Hye Won Chung , Ji Oon Lee

We study efficient algorithms for Sparse PCA in standard statistical models (spiked covariance in its Wishart form). Our goal is to achieve optimal recovery guarantees while being resilient to small perturbations. Despite a long history of…

Machine Learning · Computer Science 2020-11-13 Tommaso d'Orsi , Pravesh K. Kothari , Gleb Novikov , David Steurer

A central problem of random matrix theory is to understand the eigenvalues of spiked random matrix models, in which a prominent eigenvector is planted into a random matrix. These distributions form natural statistical models for principal…

Statistics Theory · Mathematics 2016-12-26 Amelia Perry , Alexander S. Wein , Afonso S. Bandeira , Ankur Moitra

We study the problem of graph partitioning, or clustering, in sparse networks with prior information about the clusters. Specifically, we assume that for a fraction $\rho$ of the nodes their true cluster assignments are known in advance.…

Physics and Society · Physics 2010-10-06 Armen E. Allahverdyan , Greg Ver Steeg , Aram Galstyan

Consider a two-class classification problem where the number of features is much larger than the sample size. The features are masked by Gaussian noise with mean zero and covariance matrix $\Sigma$, where the precision matrix…

Machine Learning · Statistics 2013-11-21 Yingying Fan , Jiashun Jin , Zhigang Yao

We consider the problem of Gaussian mixture clustering in the high-dimensional limit where the data consists of $m$ points in $n$ dimensions, $n,m \rightarrow \infty$ and $\alpha = m/n$ stays finite. Using exact but non-rigorous methods…

Machine Learning · Statistics 2017-03-24 Thibault Lesieur , Caterina De Bacco , Jess Banks , Florent Krzakala , Cris Moore , Lenka Zdeborová

The support recovery problem consists of determining a sparse subset of variables that is relevant in generating a set of observations. In this paper, we study the support recovery problem in the phase retrieval model consisting of noisy…

Information Theory · Computer Science 2020-09-29 Lan V. Truong , Jonathan Scarlett

A simple model to study subspace clustering is the high-dimensional $k$-Gaussian mixture model where the cluster means are sparse vectors. Here we provide an exact asymptotic characterization of the statistically optimal reconstruction…

Machine Learning · Statistics 2023-04-04 Luca Pesce , Bruno Loureiro , Florent Krzakala , Lenka Zdeborová

In the past decade, sparse principal component analysis has emerged as an archetypal problem for illustrating statistical-computational tradeoffs. This trend has largely been driven by a line of research aiming to characterize the…

Computational Complexity · Computer Science 2019-02-21 Matthew Brennan , Guy Bresler

Many problems in high-dimensional statistics appear to have a statistical-computational gap: a range of values of the signal-to-noise ratio where inference is information-theoretically possible, but (conjecturally) computationally…

Statistics Theory · Mathematics 2024-04-30 Dmitriy Kunisky , Cristopher Moore , Alexander S. Wein

In the context of sparse principal component detection, we bring evidence towards the existence of a statistical price to pay for computational efficiency. We measure the performance of a test by the smallest signal strength that it can…

Statistics Theory · Mathematics 2013-04-29 Quentin Berthet , Philippe Rigollet

A central problem of random matrix theory is to understand the eigenvalues of spiked random matrix models, introduced by Johnstone, in which a prominent eigenvector (or "spike") is planted into a random matrix. These distributions form…

Statistics Theory · Mathematics 2018-08-29 Amelia Perry , Alexander S. Wein , Afonso S. Bandeira , Ankur Moitra

In datasets where the number of parameters is fixed and the number of samples is large, principal component analysis (PCA) is a powerful dimension reduction tool. However, in many contemporary datasets, when the number of parameters is…

Probability · Mathematics 2019-02-14 Enrico Au-Yeung , Greg Zanotti

Recent work has generalized several results concerning the well-understood spiked Wigner matrix model of a low-rank signal matrix corrupted by additive i.i.d. Gaussian noise to the inhomogeneous case, where the noise has a variance profile.…

Statistics Theory · Mathematics 2025-10-10 Debsurya De , Dmitriy Kunisky