English

CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data

Computation and Language 2026-01-27 v1

Abstract

Language identification (LID) is a fundamental step in curating multilingual corpora. However, LID models still perform poorly for many languages, especially on the noisy and heterogeneous web data often used to train multilingual language models. In this paper, we introduce CommonLID, a community-driven, human-annotated LID benchmark for the web domain, covering 109 languages. Many of the included languages have been previously under-served, making CommonLID a key resource for developing more representative high-quality text corpora. We show CommonLID's value by using it, alongside five other common evaluation sets, to test eight popular LID models. We analyse our results to situate our contribution and to provide an overview of the state of the art. In particular, we highlight that existing evaluations overestimate LID accuracy for many languages in the web domain. We make CommonLID and the code used to create it available under an open, permissive license.

Keywords

Cite

@article{arxiv.2601.18026,
  title  = {CommonLID: Re-evaluating State-of-the-Art Language Identification Performance on Web Data},
  author = {Pedro Ortiz Suarez and Laurie Burchell and Catherine Arnett and Rafael Mosquera-Gómez and Sara Hincapie-Monsalve and Thom Vaughan and Damian Stewart and Malte Ostendorff and Idris Abdulmumin and Vukosi Marivate and Shamsuddeen Hassan Muhammad and Atnafu Lambebo Tonja and Hend Al-Khalifa and Nadia Ghezaiel Hammouda and Verrah Otiende and Tack Hwa Wong and Jakhongir Saydaliev and Melika Nobakhtian and Muhammad Ravi Shulthan Habibi and Chalamalasetti Kranti and Carol Muchemi and Khang Nguyen and Faisal Muhammad Adam and Luis Frentzen Salim and Reem Alqifari and Cynthia Amol and Joseph Marvin Imperial and Ilker Kesen and Ahmad Mustafid and Pavel Stepachev and Leshem Choshen and David Anugraha and Hamada Nayel and Seid Muhie Yimam and Vallerie Alexandra Putra and My Chiffon Nguyen and Azmine Toushik Wasi and Gouthami Vadithya and Rob van der Goot and Lanwenn ar C'horr and Karan Dua and Andrew Yates and Mithil Bangera and Yeshil Bangera and Hitesh Laxmichand Patel and Shu Okabe and Fenal Ashokbhai Ilasariya and Dmitry Gaynullin and Genta Indra Winata and Yiyuan Li and Juan Pablo Martínez and Amit Agarwal and Ikhlasul Akmal Hanif and Raia Abu Ahmad and Esther Adenuga and Filbert Aurelian Tjiaranata and Weerayut Buaphet and Michael Anugraha and Sowmya Vajjala and Benjamin Rice and Azril Hafizi Amirudin and Jesujoba O. Alabi and Srikant Panda and Yassine Toughrai and Bruhan Kyomuhendo and Daniel Ruffinelli and Akshata A and Manuel Goulão and Ej Zhou and Ingrid Gabriela Franco Ramirez and Cristina Aggazzotti and Konstantin Dobler and Jun Kevin and Quentin Pagès and Nicholas Andrews and Nuhu Ibrahim and Mattes Ruckdeschel and Amr Keleg and Mike Zhang and Casper Muziri and Saron Samuel and Sotaro Takeshita and Kun Kerdthaisong and Luca Foppiano and Rasul Dent and Tommaso Green and Ahmad Mustapha Wali and Kamohelo Makaaka and Vicky Feliren and Inshirah Idris and Hande Celikkanat and Abdulhamid Abubakar and Jean Maillard and Benoît Sagot and Thibault Clérice and Kenton Murray and Sarah Luger},
  journal= {arXiv preprint arXiv:2601.18026},
  year   = {2026}
}

Comments

17 pages, 7 tables, 5 figures

R2 v1 2026-07-01T09:19:29.166Z