English

A Weakly Supervised Classifier and Dataset of White Supremacist Language

Computation and Language 2023-06-29 v1

Abstract

We present a dataset and classifier for detecting the language of white supremacist extremism, a growing issue in online hate speech. Our weakly supervised classifier is trained on large datasets of text from explicitly white supremacist domains paired with neutral and anti-racist data from similar domains. We demonstrate that this approach improves generalization performance to new domains. Incorporating anti-racist texts as counterexamples to white supremacist language mitigates bias.

Keywords

Cite

@article{arxiv.2306.15732,
  title  = {A Weakly Supervised Classifier and Dataset of White Supremacist Language},
  author = {Michael Miller Yoder and Ahmad Diab and David West Brown and Kathleen M. Carley},
  journal= {arXiv preprint arXiv:2306.15732},
  year   = {2023}
}

Comments

ACL 2023 short

R2 v1 2026-06-28T11:16:03.715Z