English

Statistical Analysis of Perspective Scores on Hate Speech Detection

Computation and Language 2021-07-06 v1 Artificial Intelligence

Abstract

Hate speech detection has become a hot topic in recent years due to the exponential growth of offensive language in social media. It has proven that, state-of-the-art hate speech classifiers are efficient only when tested on the data with the same feature distribution as training data. As a consequence, model architecture plays the second role to improve the current results. In such a diverse data distribution relying on low level features is the main cause of deficiency due to natural bias in data. That's why we need to use high level features to avoid a biased judgement. In this paper, we statistically analyze the Perspective Scores and their impact on hate speech detection. We show that, different hate speech datasets are very similar when it comes to extract their Perspective Scores. Eventually, we prove that, over-sampling the Perspective Scores of a hate speech dataset can significantly improve the generalization performance when it comes to be tested on other hate speech datasets.

Keywords

Cite

@article{arxiv.2107.02024,
  title  = {Statistical Analysis of Perspective Scores on Hate Speech Detection},
  author = {Hadi Mansourifar and Dana Alsagheer and Weidong Shi and Lan Ni and Yan Huang},
  journal= {arXiv preprint arXiv:2107.02024},
  year   = {2021}
}

Comments

Accepted paper in International IJCAI Workshop on Artificial Intelligence for Social Good 2021

R2 v1 2026-06-24T03:53:58.655Z