English

Impoverished Language Technology: The Lack of (Social) Class in NLP

Computation and Language 2024-03-07 v1 Artificial Intelligence Computers and Society

Abstract

Since Labov's (1964) foundational work on the social stratification of language, linguistics has dedicated concerted efforts towards understanding the relationships between socio-demographic factors and language production and perception. Despite the large body of evidence identifying significant relationships between socio-demographic factors and language production, relatively few of these factors have been investigated in the context of NLP technology. While age and gender are well covered, Labov's initial target, socio-economic class, is largely absent. We survey the existing Natural Language Processing (NLP) literature and find that only 20 papers even mention socio-economic status. However, the majority of those papers do not engage with class beyond collecting information of annotator-demographics. Given this research lacuna, we provide a definition of class that can be operationalised by NLP researchers, and argue for including socio-economic class in future language technologies.

Keywords

Cite

@article{arxiv.2403.03874,
  title  = {Impoverished Language Technology: The Lack of (Social) Class in NLP},
  author = {Amanda Cercas Curry and Zeerak Talat and Dirk Hovy},
  journal= {arXiv preprint arXiv:2403.03874},
  year   = {2024}
}

Comments

Accepted to LREC-COLING 2024