English

Unified Sample-Optimal Property Estimation in Near-Linear Time

Machine Learning 2020-03-18 v2 Statistics Theory Machine Learning Statistics Theory

Abstract

We consider the fundamental learning problem of estimating properties of distributions over large domains. Using a novel piecewise-polynomial approximation technique, we derive the first unified methodology for constructing sample- and time-efficient estimators for all sufficiently smooth, symmetric and non-symmetric, additive properties. This technique yields near-linear-time computable estimators whose approximation values are asymptotically optimal and highly-concentrated, resulting in the first: 1) estimators achieving the O(k/(ε2logk))\mathcal{O}(k/(\varepsilon^2\log k)) min-max ε\varepsilon-error sample complexity for all kk-symbol Lipschitz properties; 2) unified near-optimal differentially private estimators for a variety of properties; 3) unified estimator achieving optimal bias and near-optimal variance for five important properties; 4) near-optimal sample-complexity estimators for several important symmetric properties over both domain sizes and confidence levels. In addition, we establish a McDiarmid's inequality under Poisson sampling, which is of independent interest.

Keywords

Cite

@article{arxiv.1911.03105,
  title  = {Unified Sample-Optimal Property Estimation in Near-Linear Time},
  author = {Yi Hao and Alon Orlitsky},
  journal= {arXiv preprint arXiv:1911.03105},
  year   = {2020}
}

Comments

Appeared at NeurIPS 2019. Fixed a few typos and minor issues in corner cases

R2 v1 2026-06-23T12:08:58.099Z