English

UnQovering Stereotyping Biases via Underspecified Questions

Computation and Language 2020-10-13 v3

Abstract

While language embeddings have been shown to have stereotyping biases, how these biases affect downstream question answering (QA) models remains unexplored. We present UNQOVER, a general framework to probe and quantify biases through underspecified questions. We show that a naive use of model scores can lead to incorrect bias estimates due to two forms of reasoning errors: positional dependence and question independence. We design a formalism that isolates the aforementioned errors. As case studies, we use this metric to analyze four important classes of stereotypes: gender, nationality, ethnicity, and religion. We probe five transformer-based QA models trained on two QA datasets, along with their underlying language models. Our broad study reveals that (1) all these models, with and without fine-tuning, have notable stereotyping biases in these classes; (2) larger models often have higher bias; and (3) the effect of fine-tuning on bias varies strongly with the dataset and the model size.

Keywords

Cite

@article{arxiv.2010.02428,
  title  = {UnQovering Stereotyping Biases via Underspecified Questions},
  author = {Tao Li and Tushar Khot and Daniel Khashabi and Ashish Sabharwal and Vivek Srikumar},
  journal= {arXiv preprint arXiv:2010.02428},
  year   = {2020}
}

Comments

Accepted at Findings of EMNLP 2020

R2 v1 2026-06-23T19:04:13.254Z