从 17 个公共网络档案收集 16K 归档网页
数字图书馆
2019-05-13 v1
摘要
我们记录了从 17 个公共网络档案构建包含 16,627 个归档网页(即 mementos)的数据集的过程,这些网页对应 3,698 个唯一实时网络 URI(统一资源标识符)。我们使用四种不同方法收集该数据集。首先,我们使用洛斯阿拉莫斯国家实验室(LANL) Memento 聚合器,从四个来源获取初始 URI 集合的 mementos:(a) Moz Top 500,(b) 我们先前研究中使用的数据集,(c) HTTP Archive,以及 (d) Web Archives for Historical Research 小组。其次,我们从已收集 mementos 的 HTML 中提取 URI,然后用这些 URI 在 LANL 聚合器中查找 mementos。第三,我们下载网络档案发布的原始页面及其关联 mementos 的 URI 列表。第四,我们从支持 Memento 协议的档案中直接请求 TimeMaps(而非通过 Memento 聚合器)收集更多 mementos。最后,由于每个档案最多 1,600 个 mementos 且能在 40 小时内下载完所有 mementos 的限制,我们将收集的 mementos 下采样至 16,627 个。
引用
@article{arxiv.1905.03836,
title = {Collecting 16K archived web pages from 17 public web archives},
author = {Mohamed Aturban and Michael L. Nelson and Michele C. Weigle and Martin Klein and Herbert Van de Sompel},
journal= {arXiv preprint arXiv:1905.03836},
year = {2019}
}
备注
21 pages