English

Translation as Augmentation: Effect of Translated Data on Assessment of Difficulty

Computation and Language 2026-07-21 v1 Machine Learning

Abstract

Reliable Text Difficulty Assessment is a prerequisite for valid text simplification workflows and personalized learning applications. However, the development of robust assessment models is severely hindered by a critical bottleneck: the scarcity of expert-annotated corpora containing fine-grained difficulty levels (e.g., CEFR), particularly for lower-resource languages. This paper addresses this data scarcity problem in the context of a low-resource European language. We propose a cross-lingual data augmentation strategy that leverages machine translation to transfer labeled resources from high-resource languages to the target low-resource language. We train BERT-based regression models to predict difficulty scores and investigate whether synthetic, translated data can effectively supplement native training sets. Our experiments demonstrate that augmenting scarce native data with machine-translated corpora significantly improves the accuracy of difficulty estimation, offering a viable solution for languages lacking extensive expert annotations.

Keywords

Cite

@article{arxiv.2607.19101,
  title  = {Translation as Augmentation: Effect of Translated Data on Assessment of Difficulty},
  author = {Yiheng Wu and Jue Hou and Roman Yangarber},
  journal= {arXiv preprint arXiv:2607.19101},
  year   = {2026}
}