English

A Theoretical Analysis of Noisy Sparse Subspace Clustering on Dimensionality-Reduced Data

Machine Learning 2019-01-24 v1 Machine Learning

Abstract

Subspace clustering is the problem of partitioning unlabeled data points into a number of clusters so that data points within one cluster lie approximately on a low-dimensional linear subspace. In many practical scenarios, the dimensionality of data points to be clustered are compressed due to constraints of measurement, computation or privacy. In this paper, we study the theoretical properties of a popular subspace clustering algorithm named sparse subspace clustering (SSC) and establish formal success conditions of SSC on dimensionality-reduced data. Our analysis applies to the most general fully deterministic model where both underlying subspaces and data points within each subspace are deterministically positioned, and also a wide range of dimensionality reduction techniques (e.g., Gaussian random projection, uniform subsampling, sketching) that fall into a subspace embedding framework (Meng & Mahoney, 2013; Avron et al., 2014). Finally, we apply our analysis to a differentially private SSC algorithm and established both privacy and utility guarantees of the proposed method.

Keywords

Cite

@article{arxiv.1610.07650,
  title  = {A Theoretical Analysis of Noisy Sparse Subspace Clustering on Dimensionality-Reduced Data},
  author = {Yining Wang and Yu-Xiang Wang and Aarti Singh},
  journal= {arXiv preprint arXiv:1610.07650},
  year   = {2019}
}

Comments

40 pages, 2 figures. A shorter version of this paper titled "A Deterministic Analysis of Noisy Sparse Subspace Clustering on Dimensionality-Reduced Data" with partial results appeared at Proceedings of the 32nd International Conference on Machine Learning (ICML) held at Lille, France in 2015

R2 v1 2026-06-22T16:30:12.067Z