English

Language corpora for the Dutch medical domain

Computation and Language 2026-04-29 v1 Artificial Intelligence

Abstract

\textbf{Background:} Dutch medical corpora are scarce, limiting NLP development. \\ \textbf{Methods:} We translated English datasets, identified medical text in generic corpora, and extracted open Dutch medical resources. \\ \textbf{Results:} The resulting corpus comprises ±\pm 35 billion tokens across the medical domain in about 100 million documents, freely available on Hugging Face. \\ \textbf{Conclusion:} This work establishes the first large-scale Dutch medical language corpus for pre-training and downstream NLP tasks.

Keywords

Cite

@article{arxiv.2604.25374,
  title  = {Language corpora for the Dutch medical domain},
  author = {B. van Es},
  journal= {arXiv preprint arXiv:2604.25374},
  year   = {2026}
}

Comments

11 pages, no figures