English

Optimal Demixing of Nonparametric Densities

Statistics Theory 2026-03-31 v1 Methodology Machine Learning Statistics Theory

Abstract

Motivated by applications in statistics and machine learning, we consider a problem of unmixing convex combinations of nonparametric densities. Suppose we observe nn groups of samples, where the iith group consists of NiN_i independent samples from a dd-variate density fi(x)=k=1Kπi(k)gk(x)f_i(x)=\sum_{k=1}^K \pi_i(k)g_k(x). Here, each gk(x)g_k(x) is a nonparametric density, and each πi\pi_i is a KK-dimensional mixed membership vector. We aim to estimate g1(x),,gK(x)g_1(x), \ldots,g_K(x). This problem generalizes topic modeling from discrete to continuous variables and finds its applications in LLMs with word embeddings. In this paper, we propose an estimator for the above problem, which modifies the classical kernel density estimator by assigning group-specific weights that are computed by topic modeling on histogram vectors and de-biased by U-statistics. For any β>0\beta>0, assuming that each gk(x)g_k(x) is in the Nikol'ski class with a smooth parameter β\beta, we show that the sum of integrated squared errors of the constructed estimators has a convergence rate that depends on nn, KK, dd, and the per-group sample size NN. We also provide a matching lower bound, which suggests that our estimator is rate-optimal.

Keywords

Cite

@article{arxiv.2603.27457,
  title  = {Optimal Demixing of Nonparametric Densities},
  author = {Jianqing Fan and Zheng Tracy Ke and Zhaoyang Shi},
  journal= {arXiv preprint arXiv:2603.27457},
  year   = {2026}
}
R2 v1 2026-07-01T11:42:34.528Z