中文

Media Cloud:开放网络上全球新闻的大规模开源采集平台

社会与信息网络 2021-05-04 v3 计算机与社会

摘要

我们给出 Media Cloud 的首个完整描述,这是一个基于超链接结构爬取、已运行超过 10 年的开源平台,对许多用途而言将是研究开放网络上媒体生态系统的最佳数据收集方式。我们记录了 Media Cloud 采集和存储哪些数据、如何处理和组织这些数据,以及其开放 API 访问与面向用户的工具背后的关键选择。我们还强调了相较于相关替代方案,Media Cloud 采集策略的优势与局限。我们概述了使用 Media Cloud 生成的两个示例数据集,并讨论了研究者如何利用该平台创建自己的数据集。

关键词

引用

@article{arxiv.2104.03702,
  title  = {Media Cloud: Massive Open Source Collection of Global News on the Open Web},
  author = {Hal Roberts and Rahul Bhargava and Linas Valiukas and Dennis Jen and Momin M. Malik and Cindy Bishop and Emily Ndulue and Aashka Dave and Justin Clark and Bruce Etling and Rob Faris and Anushka Shah and Jasmin Rubinovitz and Alexis Hope and Catherine D'Ignazio and Fernando Bermejo and Yochai Benkler and Ethan Zuckerman},
  journal= {arXiv preprint arXiv:2104.03702},
  year   = {2021}
}

备注

15 pages, 9 figures, accepted (minus the 3-page, 3-image appendix given here) for publication and forthcoming in Proceedings of the Fifteenth International AAAI Conference on Web and Social Media (ICWSM-2021)