English

Bulk Johnson-Lindenstrauss Lemmas

Probability 2023-07-18 v1 Computational Geometry Information Theory math.IT Metric Geometry Statistics Theory Statistics Theory

Abstract

For a set XX of NN points in RD\mathbb{R}^D, the Johnson-Lindenstrauss lemma provides random linear maps that approximately preserve all pairwise distances in XX -- up to multiplicative error (1±ϵ)(1\pm \epsilon) with high probability -- using a target dimension of O(ϵ2log(N))O(\epsilon^{-2}\log(N)). Certain known point sets actually require a target dimension this large -- any smaller dimension forces at least one distance to be stretched or compressed too much. What happens to the remaining distances? If we only allow a fraction η\eta of the distances to be distorted beyond tolerance (1±ϵ)(1\pm \epsilon), we show a target dimension of O(ϵ2log(4e/η)log(N)/R)O(\epsilon^{-2}\log(4e/\eta)\log(N)/R) is sufficient for the remaining distances. With the stable rank of a matrix AA as AF2/A2\lVert{A\rVert}_F^2/\lVert{A\rVert}^2, the parameter RR is the minimal stable rank over certain log(N)\log(N) sized subsets of XXX-X or their unit normalized versions, involving each point of XX exactly once. The linear maps may be taken as random matrices with i.i.d. zero-mean unit-variance sub-gaussian entries. When the data is sampled i.i.d. as a given random vector ξ\xi, refined statements are provided; the most improvement happens when ξ\xi or the unit normalized ξξ^\widehat{\xi-\xi'} is isotropic, with ξ\xi' an independent copy of ξ\xi, and includes the case of i.i.d. coordinates.

Keywords

Cite

@article{arxiv.2307.07704,
  title  = {Bulk Johnson-Lindenstrauss Lemmas},
  author = {Michael P. Casey},
  journal= {arXiv preprint arXiv:2307.07704},
  year   = {2023}
}

Comments

29 pages, no figures, the abstract is being presented at the ISI 2023 World Statistics Conference as a contributed abstract: https://www.isi2023.org/conferences/session/556/details/

R2 v1 2026-06-28T11:31:04.514Z