English

EuroLLM: Multilingual Language Models for Europe

Computation and Language 2024-09-25 v1

Abstract

The quality of open-weight LLMs has seen significant improvement, yet they remain predominantly focused on English. In this paper, we introduce the EuroLLM project, aimed at developing a suite of open-weight multilingual LLMs capable of understanding and generating text in all official European Union languages, as well as several additional relevant languages. We outline the progress made to date, detailing our data collection and filtering process, the development of scaling laws, the creation of our multilingual tokenizer, and the data mix and modeling configurations. Additionally, we release our initial models: EuroLLM-1.7B and EuroLLM-1.7B-Instruct and report their performance on multilingual general benchmarks and machine translation.

Keywords

Cite

@article{arxiv.2409.16235,
  title  = {EuroLLM: Multilingual Language Models for Europe},
  author = {Pedro Henrique Martins and Patrick Fernandes and João Alves and Nuno M. Guerreiro and Ricardo Rei and Duarte M. Alves and José Pombal and Amin Farajian and Manuel Faysse and Mateusz Klimaszewski and Pierre Colombo and Barry Haddow and José G. C. de Souza and Alexandra Birch and André F. T. Martins},
  journal= {arXiv preprint arXiv:2409.16235},
  year   = {2024}
}
R2 v1 2026-06-28T18:55:32.001Z