Rare anomalies require large datasets: About proving the existence of anomalies
Abstract
Detecting whether any anomalies exist within a dataset is crucial for effective anomaly detection, yet it remains surprisingly underexplored in anomaly detection literature. This paper presents a comprehensive study that addresses the fundamental question: When can we conclusively determine that anomalies are present? Through extensive experimentation involving over three million statistical tests across various anomaly detection tasks and algorithms, we identify a relationship between the dataset size, contamination rate, and an algorithm-dependent constant . Our results demonstrate that, for an unlabeled dataset of size and contamination rate , the condition represents a lower bound on the number of samples required to confirm anomaly existence. This threshold implies a limit to how rare anomalies can be before proving their existence becomes infeasible.
Cite
@article{arxiv.2508.09894,
title = {Rare anomalies require large datasets: About proving the existence of anomalies},
author = {Simon Klüttermann and Emmanuel Müller},
journal= {arXiv preprint arXiv:2508.09894},
year = {2025}
}
Comments
13 pages, 8 figures