English

The broken sample problem revisited: Proof of a conjecture by Bai-Hsing and high-dimensional extensions

Statistics Theory 2025-03-20 v1 Information Theory math.IT Statistics Theory

Abstract

We revisit the classical broken sample problem: Two samples of i.i.d. data points X={X1,,Xn}\mathbf{X}=\{X_1,\cdots, X_n\} and Y={Y1,,Ym}\mathbf{Y}=\{Y_1,\cdots,Y_m\} are observed without correspondence with mnm\leq n. Under the null hypothesis, X\mathbf{X} and Y\mathbf{Y} are independent. Under the alternative hypothesis, Y\mathbf{Y} is correlated with a random subsample of X\mathbf{X}, in the sense that (Xπ(i),Yi)(X_{\pi(i)},Y_i)'s are drawn independently from some bivariate distribution for some latent injection π:[m][n]\pi:[m] \to [n]. Originally introduced by DeGroot, Feder, and Goel (1971) to model matching records in census data, this problem has recently gained renewed interest due to its applications in data de-anonymization, data integration, and target tracking. Despite extensive research over the past decades, determining the precise detection threshold has remained an open problem even for equal sample sizes (m=nm=n). Assuming mm and nn grow proportionally, we show that the sharp threshold is given by a spectral and an L2L_2 condition of the likelihood ratio operator, resolving a conjecture of Bai and Hsing (2005) in the positive. These results are extended to high dimensions and settle the sharp detection thresholds for Gaussian and Bernoulli models.

Keywords

Cite

@article{arxiv.2503.14619,
  title  = {The broken sample problem revisited: Proof of a conjecture by Bai-Hsing and high-dimensional extensions},
  author = {Simiao Jiao and Yihong Wu and Jiaming Xu},
  journal= {arXiv preprint arXiv:2503.14619},
  year   = {2025}
}

Comments

35 pages, 3 figures