仇恨言论与辱骂性语言检测数据集中的种族偏见
计算与语言
2019-05-30 v1 机器学习
摘要
辱骂性语言检测技术正在被开发和应用,却很少考虑其潜在的偏见。我们考察了五个标注为仇恨言论和辱骂性语言的 Twitter 数据集中的种族偏见。我们在这些数据集上训练分类器,并比较它们对用非裔美国人英语书写的推文与用标准美国英语书写的推文的预测。结果显示所有数据集均存在系统性种族偏见的证据,因为在其上训练的分类器倾向于以显著更高的比例预测非裔美国人英语书写的推文为辱骂性。如果这些辱骂性语言检测系统投入实际使用,它们将对非裔美国社交媒体用户产生不成比例的负面影响。因此,这些系统可能歧视那些常为我们试图检测的辱骂行为之目标群体。
引用
@article{arxiv.1905.12516,
title = {Racial Bias in Hate Speech and Abusive Language Detection Datasets},
author = {Thomas Davidson and Debasmita Bhattacharya and Ingmar Weber},
journal= {arXiv preprint arXiv:1905.12516},
year = {2019}
}
备注
To appear in the proceedings of the Third Abusive Language Workshop (https://sites.google.com/view/alw3/) at the Annual Meeting for the Association for Computational Linguistics 2019. Please cite the published version