English

Gamayun's Path to Multilingual Mastery: Cost-Efficient Training of a 1.5B-Parameter LLM

Computation and Language 2025-12-30 v2

Abstract

We present Gamayun, a 1.5B-parameter multilingual language model trained entirely from scratch on 2.5T tokens. Designed for efficiency and deployment in resource-constrained environments, Gamayun addresses the lack of research on small non-English-centric LLMs by adopting a novel two-stage pre-training strategy: balanced multilingual training for cross-lingual alignment, followed by high-quality English enrichment to transfer performance gains across languages. Our model supports 12 languages, with special focus on Russian. Despite a significantly smaller training budget than comparable models, Gamayun outperforms LLaMA3.2-1B (9T tokens) on all considered benchmarks, and surpasses Qwen2.5-1.5B (18T tokens) on a wide range of English and multilingual tasks. It matches or exceeds Qwen3 (36T tokens) on most tasks outside advanced STEM, achieving state-of-the-art results in Russian, including the MERA benchmark, among the models of comparable size (1-2B parameters).

Keywords

Cite

@article{arxiv.2512.21580,
  title  = {Gamayun's Path to Multilingual Mastery: Cost-Efficient Training of a 1.5B-Parameter LLM},
  author = {Alexander Podolskiy and Semen Molokov and Timofey Gerasin and Maksim Titov and Alexey Rukhovich and Artem Khrapov and Kirill Morozov and Evgeny Tetin and Constantine Korikov and Pavel Efimov and Polina Lazukova and Yuliya Skripkar and Nikita Okhotnikov and Irina Piontkovskaya and Meng Xiaojun and Zou Xueyi and Zhang Zhenhe},
  journal= {arXiv preprint arXiv:2512.21580},
  year   = {2025}
}