English

Mind Your Inflections! Improving NLP for Non-Standard Englishes with Base-Inflection Encoding

Computation and Language 2021-05-10 v4 Artificial Intelligence Machine Learning Neural and Evolutionary Computing

Abstract

Inflectional variation is a common feature of World Englishes such as Colloquial Singapore English and African American Vernacular English. Although comprehension by human readers is usually unimpaired by non-standard inflections, current NLP systems are not yet robust. We propose Base-Inflection Encoding (BITE), a method to tokenize English text by reducing inflected words to their base forms before reinjecting the grammatical information as special symbols. Fine-tuning pretrained NLP models for downstream tasks using our encoding defends against inflectional adversaries while maintaining performance on clean data. Models using BITE generalize better to dialects with non-standard inflections without explicit training and translation models converge faster when trained with BITE. Finally, we show that our encoding improves the vocabulary efficiency of popular data-driven subword tokenizers. Since there has been no prior work on quantitatively evaluating vocabulary efficiency, we propose metrics to do so.

Keywords

Cite

@article{arxiv.2004.14870,
  title  = {Mind Your Inflections! Improving NLP for Non-Standard Englishes with Base-Inflection Encoding},
  author = {Samson Tan and Shafiq Joty and Lav R. Varshney and Min-Yen Kan},
  journal= {arXiv preprint arXiv:2004.14870},
  year   = {2021}
}

Comments

Published in the Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing

R2 v1 2026-06-23T15:12:59.467Z