English

No Text Needed: Forecasting MT Quality and Inequity from Fertility and Metadata

Computation and Language 2026-03-04 v2 Artificial Intelligence

Abstract

We show that translation quality can be predicted with surprising accuracy \textit{without ever running the translation system itself}. Using only a handful of features, token fertility ratios, token counts, and basic linguistic metadata (language family, script, and region), we can forecast ChrF scores for GPT-4o translations across 203 languages in the FLORES-200 benchmark. Gradient boosting models achieve favorable performance (R2=0.66R^{2}=0.66 for XX\rightarrowEnglish and R2=0.72R^{2}=0.72 for English\rightarrowXX). Feature importance analyses reveal that typological factors dominate predictions into English, while fertility plays a larger role for translations into diverse target languages. These findings suggest that translation quality is shaped by both token-level fertility and broader linguistic typology, offering new insights for multilingual evaluation and quality estimation.

Keywords

Cite

@article{arxiv.2509.05425,
  title  = {No Text Needed: Forecasting MT Quality and Inequity from Fertility and Metadata},
  author = {Jessica M. Lundin and Ada Zhang and David Adelani and Cody Carroll},
  journal= {arXiv preprint arXiv:2509.05425},
  year   = {2026}
}