English

Sociodemographic Bias in Language Models: A Survey and Forward Path

Computation and Language 2024-08-15 v5 Artificial Intelligence Machine Learning

Abstract

Sociodemographic bias in language models (LMs) has the potential for harm when deployed in real-world settings. This paper presents a comprehensive survey of the past decade of research on sociodemographic bias in LMs, organized into a typology that facilitates examining the different aims: types of bias, quantifying bias, and debiasing techniques. We track the evolution of the latter two questions, then identify current trends and their limitations, as well as emerging techniques. To guide future research towards more effective and reliable solutions, and to help authors situate their work within this broad landscape, we conclude with a checklist of open questions.

Keywords

Cite

@article{arxiv.2306.08158,
  title  = {Sociodemographic Bias in Language Models: A Survey and Forward Path},
  author = {Vipul Gupta and Pranav Narayanan Venkit and Shomir Wilson and Rebecca J. Passonneau},
  journal= {arXiv preprint arXiv:2306.08158},
  year   = {2024}
}

Comments

23 pages, 3 figure

R2 v1 2026-06-28T11:04:30.872Z