English

How Should We Model the Probability of a Language?

Computation and Language 2026-02-10 v1

Abstract

Of the over 7,000 languages spoken in the world, commercial language identification (LID) systems only reliably identify a few hundred in written form. Research-grade systems extend this coverage under certain circumstances, but for most languages coverage remains patchy or nonexistent. This position paper argues that this situation is largely self-imposed. In particular, it arises from a persistent framing of LID as decontextualized text classification, which obscures the central role of prior probability estimation and is reinforced by institutional incentives that favor global, fixed-prior models. We argue that improving coverage for tail languages requires rethinking LID as a routing problem and developing principled ways to incorporate environmental cues that make languages locally plausible.

Keywords

Cite

@article{arxiv.2602.08951,
  title  = {How Should We Model the Probability of a Language?},
  author = {Rasul Dent and Pedro Ortiz Suarez and Thibault Clérice and Benoît Sagot},
  journal= {arXiv preprint arXiv:2602.08951},
  year   = {2026}
}

Comments

Accepted for Vardial 2026

R2 v1 2026-07-01T10:28:25.534Z