English

Compressibility Barriers to Neighborhood-Preserving Data Visualizations

Computational Geometry 2026-01-19 v2 Metric Geometry

Abstract

To what extent is it possible to visualize high-dimensional data in two- or three-dimensional plots? We reframe this question in terms of embedding nn-vertex graphs (representing the neighborhood structure of the input points) into metric spaces of low doubling dimension dd in such a way that keeps neighbors close and non-neighbors far. This notion of neighbor preservation can be understood as a considerably weaker embedding constraint than near-isometry, yet it is similarly as demanding in terms of how the minimum required dimension scales with the number of points. We show that for an overwhelming fraction of graphs, d=Θ(logn)d = \Theta(\log n) is both necessary and sufficient for neighbor preservation. Even sparse regular graphs, which represent more restricted neighborhood connectivity structures, typically require d=Ω(logn/loglogn)d= \Omega(\log n / \log\log n). The landscape changes dramatically when embedding into normed spaces: general graphs become exponentially harder to embed, requiring d=Ω(n)d=\Omega(n), while sparse regular graphs continue to admit d=O(logn)d = O(\log n). Finally, we study the implications of these results for visualizing data with intrinsic cluster structure. We show that graphs produced from a planted partition model with kk clusters on nn points typically require d=Ω(logn)d=\Omega(\log n), even when the cluster structure is salient. These results challenge the aspiration that constant-dimensional visualizations can faithfully preserve neighborhood structure.

Keywords

Cite

@article{arxiv.2508.07119,
  title  = {Compressibility Barriers to Neighborhood-Preserving Data Visualizations},
  author = {Szymon Snoeck and Noah Bergam and Nakul Verma},
  journal= {arXiv preprint arXiv:2508.07119},
  year   = {2026}
}
R2 v1 2026-07-01T04:42:43.398Z