English

Intriguing Properties of Compression on Multilingual Models

Computation and Language 2022-11-29 v2 Artificial Intelligence

Abstract

Multilingual models are often particularly dependent on scaling to generalize to a growing number of languages. Compression techniques are widely relied upon to reconcile the growth in model size with real world resource constraints, but compression can have a disparate effect on model performance for low-resource languages. It is thus crucial to understand the trade-offs between scale, multilingualism, and compression. In this work, we propose an experimental framework to characterize the impact of sparsifying multilingual pre-trained language models during fine-tuning. Applying this framework to mBERT named entity recognition models across 40 languages, we find that compression confers several intriguing and previously unknown generalization properties. In contrast to prior findings, we find that compression may improve model robustness over dense models. We additionally observe that under certain sparsification regimes compression may aid, rather than disproportionately impact the performance of low-resource languages.

Keywords

Cite

@article{arxiv.2211.02738,
  title  = {Intriguing Properties of Compression on Multilingual Models},
  author = {Kelechi Ogueji and Orevaoghene Ahia and Gbemileke Onilude and Sebastian Gehrmann and Sara Hooker and Julia Kreutzer},
  journal= {arXiv preprint arXiv:2211.02738},
  year   = {2022}
}

Comments

Accepted to EMNLP 2022

R2 v1 2026-06-28T05:13:42.735Z