English

FuLG: 150B Romanian Corpus for Language Model Pretraining

Computation and Language 2024-07-19 v1

Abstract

Research in the field of language models is rapidly evolving, with many open models being released to the public. Openly available pretraining corpora usually focus on only a handful of languages, with many others either missing completely or extremely underrepresented. In this report, we introduce FuLG, a hundred-fifty-billion-token Romanian corpus extracted from CommonCrawl. We present our methodology for filtering FuLG and compare it via ablation studies against existing Romanian corpora.

Keywords

Cite

@article{arxiv.2407.13657,
  title  = {FuLG: 150B Romanian Corpus for Language Model Pretraining},
  author = {Vlad-Andrei Bădoiu and Mihai-Valentin Dumitru and Alexandru M. Gherghescu and Alexandru Agache and Costin Raiciu},
  journal= {arXiv preprint arXiv:2407.13657},
  year   = {2024}
}
R2 v1 2026-06-28T17:46:15.475Z