English

PCA of probability measures: Sparse and Dense sampling regimes

Machine Learning 2026-02-03 v1 Machine Learning

Abstract

A common approach to perform PCA on probability measures is to embed them into a Hilbert space where standard functional PCA techniques apply. While convergence rates for estimating the embedding of a single measure from mm samples are well understood, the literature has not addressed the setting involving multiple measures. In this paper, we study PCA in a double asymptotic regime where nn probability measures are observed, each through mm samples. We derive convergence rates of the form n1/2+mαn^{-1/2} + m^{-\alpha} for the empirical covariance operator and the PCA excess risk, where α>0\alpha>0 depends on the chosen embedding. This characterizes the relationship between the number nn of measures and the number mm of samples per measure, revealing a sparse (small mm) to dense (large mm) transition in the convergence behavior. Moreover, we prove that the dense-regime rate is minimax optimal for the empirical covariance error. Our numerical experiments validate these theoretical rates and demonstrate that appropriate subsampling preserves PCA accuracy while reducing computational cost.

Keywords

Cite

@article{arxiv.2602.02190,
  title  = {PCA of probability measures: Sparse and Dense sampling regimes},
  author = {Gachon Erell and Jérémie Bigot and Elsa Cazelles},
  journal= {arXiv preprint arXiv:2602.02190},
  year   = {2026}
}
R2 v1 2026-07-01T09:32:00.389Z