English

Optimal Algorithms and Lower Bounds for Testing Closeness of Structured Distributions

Data Structures and Algorithms 2015-08-25 v1 Information Theory math.IT Statistics Theory Statistics Theory

Abstract

We give a general unified method that can be used for L1L_1 {\em closeness testing} of a wide range of univariate structured distribution families. More specifically, we design a sample optimal and computationally efficient algorithm for testing the equivalence of two unknown (potentially arbitrary) univariate distributions under the Ak\mathcal{A}_k-distance metric: Given sample access to distributions with density functions p,q:IRp, q: I \to \mathbb{R}, we want to distinguish between the cases that p=qp=q and pqAkϵ\|p-q\|_{\mathcal{A}_k} \ge \epsilon with probability at least 2/32/3. We show that for any k2,ϵ>0k \ge 2, \epsilon>0, the {\em optimal} sample complexity of the Ak\mathcal{A}_k-closeness testing problem is Θ(max{k4/5/ϵ6/5,k1/2/ϵ2})\Theta(\max\{ k^{4/5}/\epsilon^{6/5}, k^{1/2}/\epsilon^2 \}). This is the first o(k)o(k) sample algorithm for this problem, and yields new, simple L1L_1 closeness testers, in most cases with optimal sample complexity, for broad classes of structured distributions.

Keywords

Cite

@article{arxiv.1508.05538,
  title  = {Optimal Algorithms and Lower Bounds for Testing Closeness of Structured Distributions},
  author = {Ilias Diakonikolas and Daniel M. Kane and Vladimir Nikishkin},
  journal= {arXiv preprint arXiv:1508.05538},
  year   = {2015}
}

Comments

27 pages, to appear in FOCS'15

R2 v1 2026-06-22T10:39:29.873Z