The Danish Gigaword Project
Computation and Language
2021-05-14 v3
Abstract
Danish language technology has been hindered by a lack of broad-coverage corpora at the scale modern NLP prefers. This paper describes the Danish Gigaword Corpus, the result of a focused effort to provide a diverse and freely-available one billion word corpus of Danish text. The Danish Gigaword corpus covers a wide array of time periods, domains, speakers' socio-economic status, and Danish dialects.
Keywords
Cite
@article{arxiv.2005.03521,
title = {The Danish Gigaword Project},
author = {Leon Strømberg-Derczynski and Manuel R. Ciosici and Rebekah Baglini and Morten H. Christiansen and Jacob Aarup Dalsgaard and Riccardo Fusaroli and Peter Juel Henrichsen and Rasmus Hvingelby and Andreas Kirkedal and Alex Speed Kjeldsen and Claus Ladefoged and Finn Årup Nielsen and Malte Lau Petersen and Jonathan Hvithamar Rystrøm and Daniel Varab},
journal= {arXiv preprint arXiv:2005.03521},
year = {2021}
}
Comments
Identical to the NoDaLiDa 2021 version