English

DiRe-RAPIDS: Topology-faithful dimensionality reduction at scale

Machine Learning 2026-04-30 v2 Artificial Intelligence Software Engineering Social and Information Networks

Abstract

Dimensionality reduction methods such as UMAP and t-SNE are central tools for visualising high-dimensional data, but their local-neighborhood objectives can preserve sampling noise while distorting global topology. We show that standard local metrics reward this noise memorisation: top-performing embeddings invent cycles and disconnected islands absent from the data. We introduce a topology-faithfulness benchmark based on noisy manifolds with known homology, tune DiRe against it, and find Pareto-optimal configurations that match or beat GPU-accelerated UMAP on classification while recovering exact first Betti numbers on stress tests. On 723K arXiv paper embeddings, DiRe preserves 3-4 times more topological structure than UMAP at comparable wall-clock.

Keywords

Cite

@article{arxiv.2604.25209,
  title  = {DiRe-RAPIDS: Topology-faithful dimensionality reduction at scale},
  author = {Alexander Kolpakov and Igor Rivin},
  journal= {arXiv preprint arXiv:2604.25209},
  year   = {2026}
}

Comments

5 pages, 4 figures, fixed broken URLs in comments; GitHub repositories https://github.com/sashakolpakov/dire-rapids | https://github.com/igorrivin/dire-rapids-arxiv | HuggingFace dataset https://huggingface.co/datasets/igriv/dire-arxiv-bge-small-embeddings