English

Sparse Principal Component Analysis via Wavelets for Distributed Data

Methodology 2026-08-05 v1

Abstract

The large volume of data and concerns about data privacy have motivated the development of techniques for distributed data, a problem also known as federated learning. In this scenario, sub-samples of the data are divided across different machines, and statistics must be computed over that data without direct access to the full sample. Johnstone & Lu (2009, JASA) show that principal component analysis (PCA) is statistically inconsistent in the high-dimensional regime, and propose a way to recover consistency through wavelet-based sparsification and variable selection. Fan et al. (2019, AoS) show a way to perform this same estimation -- specifically, to estimate the eigenspace that would be obtained if all the data were pooled together, even though it remains effectively distributed -- without addressing the high-dimensional regime. This work incorporates the wavelet-based sparsification of Johnstone & Lu (2009) into the distributed PCA framework of Fan et al. (2019), aiming to reduce communication cost without compromising the quality of the eigenspace estimation. Simulations across d[52,5000]d \in [52, 5000] show that the proposed method overtakes Fan et al. (2019) in estimation error beyond a clear dimensional threshold (d152d \geq 152 for λ=25\lambda=25, d252d \geq 252 for λ=50\lambda=50), while transmitting systematically fewer coefficients throughout the entire range studied. This study was financed by the Sao Paulo Research Foundation (FAPESP), Brazil. Process Number #2023/02538-0 and Number #2025/21250-2.

Cite

@article{arxiv.2608.05386,
  title  = {Sparse Principal Component Analysis via Wavelets for Distributed Data},
  author = {Giovanni Barbosa Herrero and Rodney Vasconcelos Fonseca and Aluísio Pinheiro},
  journal= {arXiv preprint arXiv:2608.05386},
  year   = {2026}
}

Comments

10 pages, 3 figures, 1 table