English

On the Limits of Model Merging for Multilinguality in Pre-Training

Computation and Language 2026-05-26 v1

Abstract

Endowing models with consistent multilingual performance can be achieved by mixing pre-training data, or post-training approaches such as language-specific model merging. In this work, we test whether merging can be applied to monolingually pre-trained models. We conduct a controlled study on the efficacy of mixed, merged, and monolingual pre-training setups. We find that while monolingual pre-training results in strong in-language performance, merging any combination of monolingual models leads to performance collapse due to interference. Our analysis suggests representational similarity is a prerequisite for model merging. We therefore conclude that the flexibility of merging in fine-tuning does not extend trivially to language-specific pre-training.

Keywords

Cite

@article{arxiv.2605.25846,
  title  = {On the Limits of Model Merging for Multilinguality in Pre-Training},
  author = {Seth Aycock and Fedor Vitiugin and Aleksandr Umnov and Christof Monz and Khalil Sima'an},
  journal= {arXiv preprint arXiv:2605.25846},
  year   = {2026}
}

Comments

MeLLM Workshop 2026

R2 v1 2026-07-22T07:32:32.108Z