Computation and Language · Computer Science
CC-GPX: Extracting High-Quality Annotated Geospatial Data from Common Crawl
Ilya Ilyankou, Meihui Wang, Stefano Cavazzi, James Haworth
2026-05-07
Computation and Language · Computer Science
Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew +4
2021-10-01
Computation and Language · Computer Science
Building a Web-Scale Dependency-Parsed Corpus from CommonCrawl
Alexander Panchenko, Eugen Ruppert, Stefano Faralli, Simone Paolo Ponzetto +1
2018-03-01
Computation and Language · Computer Science
Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets
Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab +48
2022-02-22
Computation and Language · Computer Science
CCpdf: Building a High Quality Corpus for Visually Rich Documents from Web Crawl Data
Michał Turski, Tomasz Stanisławek, Karol Kaczmarek, Paweł Dyda +1
2023-06-07
Computation and Language · Computer Science
Quantifying Geospatial in the Common Crawl Corpus
Ilya Ilyankou, Meihui Wang, Stefano Cavazzi, James Haworth
2026-05-07
Information Retrieval · Computer Science
Resource Selection for Federated Search on the Web
Dong Nguyen, Thomas Demeester, Dolf Trieschnigg, Djoerd Hiemstra
2016-09-16
Computation and Language · Computer Science
CCAligned: A Massive Collection of Cross-Lingual Web-Document Pairs
Ahmed El-Kishky, Vishrav Chaudhary, Francisco Guzman, Philipp Koehn
2020-10-13
Networking and Internet Architecture · Computer Science
Decoding the structure of the WWW: facts versus sampling biases
M. Angeles Serrano, Ana Maguitman, Marian Boguna, Santo Fortunato +1
2008-01-23
Computer Vision and Pattern Recognition · Computer Science
A Large Visual, Qualitative and Quantitative Dataset of Web Pages
Christian Mejia-Escobar, Miguel Cazorla, Ester Martinez-Martin
2021-05-18
Computation and Language · Computer Science
The Growing Gains and Pains of Iterative Web Corpora Crawling: Insights from South Slavic CLASSLA-web 2.0 Corpora
Taja Kuzman Pungeršek, Peter Rupnik, Vít Suchomel, Nikola Ljubešić
2026-03-02
Information Retrieval · Computer Science
OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents
Hugo Laurençon, Lucile Saulnier, Léo Tronchon, Stas Bekman +8
2023-08-22
Computation and Language · Computer Science
The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru +5
2023-06-05
Computation and Language · Computer Science
esCorpius: A Massive Spanish Crawling Corpus
Asier Gutiérrez-Fandiño, David Pérez-Fernández, Jordi Armengol-Estapé, David Griol +1
2022-07-04
Computation and Language · Computer Science
Going over Fine Web with a Fine-Tooth Comb: Technical Report of Indexing Fine Web for Problematic Content Search and Retrieval
Inés Altemir Marinas, Anastasiia Kucherenko, Andrei Kucharavy
2025-09-01
Information Retrieval · Computer Science
ClueWeb22: 10 Billion Web Documents with Visual and Semantic Information
Arnold Overwijk, Chenyan Xiong, Xiao Liu, Cameron VandenBerg +1
2022-12-05