English

High-Resource Methodological Bias in Low-Resource Investigations

Computation and Language 2022-11-15 v1

Abstract

The central bottleneck for low-resource NLP is typically regarded to be the quantity of accessible data, overlooking the contribution of data quality. This is particularly seen in the development and evaluation of low-resource systems via down sampling of high-resource language data. In this work we investigate the validity of this approach, and we specifically focus on two well-known NLP tasks for our empirical investigations: POS-tagging and machine translation. We show that down sampling from a high-resource language results in datasets with different properties than the low-resource datasets, impacting the model performance for both POS-tagging and machine translation. Based on these results we conclude that naive down sampling of datasets results in a biased view of how well these systems work in a low-resource scenario.

Keywords

Cite

@article{arxiv.2211.07534,
  title  = {High-Resource Methodological Bias in Low-Resource Investigations},
  author = {Maartje ter Hoeve and David Grangier and Natalie Schluter},
  journal= {arXiv preprint arXiv:2211.07534},
  year   = {2022}
}
R2 v1 2026-06-28T05:49:38.478Z