Measuring What the Crawler Sees: Discovery Curves, Core Persistence, and Shell Dynamics in Longitudinal Web Crawls
Abstract
A longitudinal web crawl is a sequence of partial samples of an evolving URL population. Pairwise containment between two crawls is the standard probe; under a simple \emph{urn} model of the crawl -- each round samples a fraction of the URLs and replaces a fraction -- it recovers two interpretable rates, per-round survival and coverage , but treats the population as uniform and consumes one pair at a time. In this work, we define a formal language for talking about a crawl. We extend this analysis with the \emph{discovery curve} , the cumulative URL footprint over a sliding window of crawls starting at , which under the same urn model is also a closed-form function of . Containment and the discovery curve are then two projections of one process: independent fits agree on when the urn is homogeneous, so any disagreement is itself a measurement. Applied to Common Crawl (2020--2025, domain granularity) and to the German Academic Web (GAW, URL granularity), the two projections disagree on both archives, and a two-component urn with a persistent core fraction alongside shell parameters reconciles the disagreement. A residual on remains, signaling that the shell itself is not homogeneous; is recorded as the scalar entry point to a rank-resolved generalization, which is left to follow-up work. \keywords{web archive \and crawl coverage \and discovery curve \and urn model \and two-component model \and URL lifetime}
Cite
@article{arxiv.2607.13636,
title = {Measuring What the Crawler Sees: Discovery Curves, Core Persistence, and Shell Dynamics in Longitudinal Web Crawls},
author = {Michael Paris and Hande Celikkanat and Luca Foppiano},
journal= {arXiv preprint arXiv:2607.13636},
year = {2026}
}
Comments
16 pages, 4 figures, web metrics