English

Efficient estimation of the cardinality of large data sets

Statistics Theory 2011-04-25 v3 Probability Statistics Theory

Abstract

F.Giroire has recently proposed an algorithm which returns the approximate number of distincts elements in a large sequence of words, under strong constraints coming from the analysis of large data bases. His estimation is based on statistical properties of uniform random variables in [0,1][0,1]. In this note we propose an optimal estimation, using Kullback information and estimation theory.

Cite

@article{arxiv.math/0701347,
  title  = {Efficient estimation of the cardinality of large data sets},
  author = {Philippe Chassaing and Lucas Gerin},
  journal= {arXiv preprint arXiv:math/0701347},
  year   = {2011}
}

Comments

Extended and improved version of the published paper

R2 v1 2026-07-22T17:49:15.687Z