Information Theoretic Limits of Cardinality Estimation: Fisher Meets Shannon
Abstract
Estimating the cardinality (number of distinct elements) of a large multiset is a classic problem in streaming and sketching. In this paper we study the intrinsic tradeoff between the space complexity of the sketch and its estimation error. We define a new measure of efficiency for data sketches called the Fisher-Shannon (FiSh) number . It captures the tension between the limiting Shannon entropy () of the sketch and its normalized Fisher information () that characterizes the variance of a statistically efficient, asymptotically unbiased estimator. Our aim in introducing the FiSh-number is to build the mathematical machinery necessary to argue for precise optimality, rather than asymptotic optimality, up to large constant factors. Our results are as follows. [1] We prove that all base- variants of Flajolet and Martin's PCSA sketch have FiSh-number and that every base- variant of HyperLogLog has FiSh-number worse than , but that they tend to in the limit as . Here are precisely defined constants. [2] We describe a sketch called Fishmonger that is based on a smoothed, entropy-compressed variant of PCSA with a different estimator function. Fishmonger processes a multiset of such that at all times, w.h.p., its space is bits and its standard error is . For example, to achieve a 1% standard error, one needs a little more than 19,800 bits, or kilobytes. [3] Finally, we give circumstantial evidence that is the optimum FiSh-number of mergeable sketches for Cardinality Estimation. We define a natural subset of mergeable sketches called linearizable sketches and prove that no member of this class can beat . The popular mergeable sketches are, in fact, also linearizable.
Keywords
Cite
@article{arxiv.2007.08051,
title = {Information Theoretic Limits of Cardinality Estimation: Fisher Meets Shannon},
author = {Seth Pettie and Dingyu Wang},
journal= {arXiv preprint arXiv:2007.08051},
year = {2026}
}