English

Small Languages, Big Models: A Study of Continual Training on Languages of Norway

Computation and Language 2025-02-04 v2

Abstract

Training large language models requires vast amounts of data, posing a challenge for less widely spoken languages like Norwegian and even more so for truly low-resource languages like Northern S\'ami. To address this issue, we present a novel three-stage continual training approach that substantially improves the downstream performance together with the inference efficiency for the target languages. Based on our findings, we train, evaluate, and openly release a new generative language model for Norwegian Bokm\r{a}l, Nynorsk, and Northern S\'ami with 11.4 billion parameters: NorMistral-11B.

Keywords

Cite

@article{arxiv.2412.06484,
  title  = {Small Languages, Big Models: A Study of Continual Training on Languages of Norway},
  author = {David Samuel and Vladislav Mikhailov and Erik Velldal and Lilja Øvrelid and Lucas Georges Gabriel Charpentier and Andrey Kutuzov and Stephan Oepen},
  journal= {arXiv preprint arXiv:2412.06484},
  year   = {2025}
}

Comments

Published at NoDaLiDa 2025