English

The Danish Gigaword Project

Computation and Language 2021-05-14 v3

Abstract

Danish language technology has been hindered by a lack of broad-coverage corpora at the scale modern NLP prefers. This paper describes the Danish Gigaword Corpus, the result of a focused effort to provide a diverse and freely-available one billion word corpus of Danish text. The Danish Gigaword corpus covers a wide array of time periods, domains, speakers' socio-economic status, and Danish dialects.

Keywords

Cite

@article{arxiv.2005.03521,
  title  = {The Danish Gigaword Project},
  author = {Leon Strømberg-Derczynski and Manuel R. Ciosici and Rebekah Baglini and Morten H. Christiansen and Jacob Aarup Dalsgaard and Riccardo Fusaroli and Peter Juel Henrichsen and Rasmus Hvingelby and Andreas Kirkedal and Alex Speed Kjeldsen and Claus Ladefoged and Finn Årup Nielsen and Malte Lau Petersen and Jonathan Hvithamar Rystrøm and Daniel Varab},
  journal= {arXiv preprint arXiv:2005.03521},
  year   = {2021}
}

Comments

Identical to the NoDaLiDa 2021 version

R2 v1 2026-06-23T15:23:04.970Z