English

High-Dimensional Procrustes Matching via Tree Counts

Machine Learning 2026-07-09 v1 Information Theory Machine Learning Statistics Theory

Abstract

Suppose we observe two sets of nn Gaussian vectors in Rd\mathbb{R}^d, with the promise that, after applying a permutation of [n][n] and a rotation of Rd\mathbb{R}^d, the two sets are ρ\rho-correlated. The Procrustes matching problem asks us to recover the unknown permutation of [n][n] that aligns the two sets. The problem is well-studied in the low-dimensional regime d=O(logn)d=O(\log n), but the high-dimensional regime dlognd\gg \log n has remained largely uncharted: prior matching guarantees require nearly perfect correlation ρ=1o(1)\rho=1-o(1), even for information-theoretic recovery. Our main result is a polynomial-time algorithm for exact recovery at constant correlation. The algorithm works by computing and comparing weighted counts of a specially chosen family of ``wide'' trees. So long as dpolylog(n)d\ge \mathrm{polylog}(n), the algorithm succeeds with high probability for any ρ2>α\rho^2>\sqrt{\alpha}, where α0.338\alpha\approx 0.338 is Otter's tree-counting constant. We complement this algorithmic result with an improved information-theoretic guarantee, showing that exact recovery is possible when ρ2max{logn/d,logn/n}\rho^2 \gtrsim \max\{\log n/d,\sqrt{\log n/n}\}. We also carry out a low-degree advantage calculation, which suggests that the condition ρ2>α\rho^2 > \sqrt{\alpha} is necessary for any tree-counting algorithm.

Cite

@article{arxiv.2607.08538,
  title  = {High-Dimensional Procrustes Matching via Tree Counts},
  author = {Xiaochun Niu and Tselil Schramm and Jiaming Xu},
  journal= {arXiv preprint arXiv:2607.08538},
  year   = {2026}
}