Sparse Principal Component Analysis via Wavelets for Distributed Data
Abstract
The large volume of data and concerns about data privacy have motivated the development of techniques for distributed data, a problem also known as federated learning. In this scenario, sub-samples of the data are divided across different machines, and statistics must be computed over that data without direct access to the full sample. Johnstone & Lu (2009, JASA) show that principal component analysis (PCA) is statistically inconsistent in the high-dimensional regime, and propose a way to recover consistency through wavelet-based sparsification and variable selection. Fan et al. (2019, AoS) show a way to perform this same estimation -- specifically, to estimate the eigenspace that would be obtained if all the data were pooled together, even though it remains effectively distributed -- without addressing the high-dimensional regime. This work incorporates the wavelet-based sparsification of Johnstone & Lu (2009) into the distributed PCA framework of Fan et al. (2019), aiming to reduce communication cost without compromising the quality of the eigenspace estimation. Simulations across show that the proposed method overtakes Fan et al. (2019) in estimation error beyond a clear dimensional threshold ( for , for ), while transmitting systematically fewer coefficients throughout the entire range studied. This study was financed by the Sao Paulo Research Foundation (FAPESP), Brazil. Process Number #2023/02538-0 and Number #2025/21250-2.
Cite
@article{arxiv.2608.05386,
title = {Sparse Principal Component Analysis via Wavelets for Distributed Data},
author = {Giovanni Barbosa Herrero and Rodney Vasconcelos Fonseca and Aluísio Pinheiro},
journal= {arXiv preprint arXiv:2608.05386},
year = {2026}
}
Comments
10 pages, 3 figures, 1 table