English

MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs

Computation and Language 2026-05-01 v4

Abstract

We introduce MultiBLiMP 1.0, a massively multilingual benchmark of linguistic minimal pairs, covering 101 languages and 2 types of subject-verb agreement, containing more than 128,000 minimal pairs. Our minimal pairs are created using a fully automated pipeline, leveraging the large-scale linguistic resources of Universal Dependencies and UniMorph. MultiBLiMP 1.0 evaluates abilities of LLMs at an unprecedented multilingual scale, and highlights the shortcomings of the current state-of-the-art in modelling low-resource languages.

Keywords

Cite

@article{arxiv.2504.02768,
  title  = {MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs},
  author = {Jaap Jumelet and Leonie Weissweiler and Joakim Nivre and Arianna Bisazza},
  journal= {arXiv preprint arXiv:2504.02768},
  year   = {2026}
}

Comments

Published in TACL, MIT Press