你只需要“Leet”:规避仇恨言论检测 AI
密码学与安全
2025-05-23 v1 计算与语言
机器学习
摘要
社交媒体和在线论坛正日益流行,但这些平台也被用于传播仇恨言论。本文设计了黑盒技术,通过生成能够欺骗基于深度学习的仇恨言论检测模型的扰动来保护用户免受仇恨言论的侵害,从而降低这些模型的效率。我们还确保对仇恨言论原始意义的最小化改变。我们最佳的扰动攻击成功能够规避 86.8% 的仇恨文本的仇恨言论检测。
引用
@article{arxiv.2505.16263,
title = {All You Need is "Leet": Evading Hate-speech Detection AI},
author = {Sampanna Yashwant Kahu and Naman Ahuja},
journal= {arXiv preprint arXiv:2505.16263},
year = {2025}
}
备注
10 pages, 22 figures, The source code and data used in this work is available at: https://github.com/SampannaKahu/all_you_need_is_leet