English

All You Need is "Leet": Evading Hate-speech Detection AI

Cryptography and Security 2025-05-23 v1 Computation and Language Machine Learning

Abstract

Social media and online forums are increasingly becoming popular. Unfortunately, these platforms are being used for spreading hate speech. In this paper, we design black-box techniques to protect users from hate-speech on online platforms by generating perturbations that can fool state of the art deep learning based hate speech detection models thereby decreasing their efficiency. We also ensure a minimal change in the original meaning of hate-speech. Our best perturbation attack is successfully able to evade hate-speech detection for 86.8 % of hateful text.

Keywords

Cite

@article{arxiv.2505.16263,
  title  = {All You Need is "Leet": Evading Hate-speech Detection AI},
  author = {Sampanna Yashwant Kahu and Naman Ahuja},
  journal= {arXiv preprint arXiv:2505.16263},
  year   = {2025}
}

Comments

10 pages, 22 figures, The source code and data used in this work is available at: https://github.com/SampannaKahu/all_you_need_is_leet

R2 v1 2026-07-01T02:30:31.658Z