English

Structure preservation via the Wasserstein distance

Statistics Theory 2025-01-13 v3 Functional Analysis Probability Statistics Theory

Abstract

We show that under minimal assumptions on a random vector XRdX\in\mathbb{R}^d and with high probability, given mm independent copies of XX, the coordinate distribution of each vector (Xi,θ)i=1m(\langle X_i,\theta \rangle)_{i=1}^m is dictated by the distribution of the true marginal X,θ\langle X,\theta \rangle. Specifically, we show that with high probability, supθSd1(1mi=1mXi,θλiθ2)1/2c(dm)1/4,\sup_{\theta \in S^{d-1}} \left( \frac{1}{m}\sum_{i=1}^m \left|\langle X_i,\theta \rangle^\sharp - \lambda^\theta_i \right|^2 \right)^{1/2} \leq c \left( \frac{d}{m} \right)^{1/4}, where λiθ=m(i1m,im]FX,θ1(u)du\lambda^{\theta}_i = m\int_{(\frac{i-1}{m}, \frac{i}{m}]} F_{ \langle X,\theta \rangle }^{-1}(u)\,du and aa^\sharp denotes the monotone non-decreasing rearrangement of aa. Moreover, this estimate is optimal. The proof follows from a sharp estimate on the worst Wasserstein distance between a marginal of XX and its empirical counterpart, 1mi=1mδXi,θ\frac{1}{m} \sum_{i=1}^m \delta_{\langle X_i, \theta \rangle}.

Keywords

Cite

@article{arxiv.2209.07058,
  title  = {Structure preservation via the Wasserstein distance},
  author = {Daniel Bartl and Shahar Mendelson},
  journal= {arXiv preprint arXiv:2209.07058},
  year   = {2025}
}

Comments

Original paper [v1] was split into two papers. Here is the first part. Second part is now called "Optimal non-gaussian Dvoretzky-Milman embeddings"

R2 v1 2026-06-28T01:20:14.268Z