English

Pretraining Finnish ModernBERTs

Computation and Language 2025-11-13 v1

Abstract

This paper reports on pretraining ModernBERT encoder models in six different sizes, ranging from 51M to 475M parameters, with a focus on limited multilingualism, emphasizing languages relevant to Finland. Our models are competitive with, or superior to, existing multilingual models. They outperform monolingual models on tasks that require a context longer than 512 tokens. We present empirical results on using different data in the final stage of training. The code and models are publicly released.

Keywords

Cite

@article{arxiv.2511.09213,
  title  = {Pretraining Finnish ModernBERTs},
  author = {Akseli Reunamo and Laura-Maria Peltonen and Hans Moen and Sampo Pyysalo},
  journal= {arXiv preprint arXiv:2511.09213},
  year   = {2025}
}
R2 v1 2026-07-01T07:33:46.193Z