English

The first open machine translation system for the Chechen language

Computation and Language 2025-07-18 v1

Abstract

We introduce the first open-source model for translation between the vulnerable Chechen language and Russian, and the dataset collected to train and evaluate it. We explore fine-tuning capabilities for including a new language into a large language model system for multilingual translation NLLB-200. The BLEU / ChrF++ scores for our model are 8.34 / 34.69 and 20.89 / 44.55 for translation from Russian to Chechen and reverse direction, respectively. The release of the translation models is accompanied by the distribution of parallel words, phrases and sentences corpora and multilingual sentence encoder adapted to the Chechen language.

Keywords

Cite

@article{arxiv.2507.12672,
  title  = {The first open machine translation system for the Chechen language},
  author = {Abu-Viskhan A. Umishov and Vladislav A. Grigorian},
  journal= {arXiv preprint arXiv:2507.12672},
  year   = {2025}
}

Comments

7 pages

R2 v1 2026-07-01T04:05:12.117Z