破解密码:通过编码分类提高隐式仇恨言论检测
计算与语言
2025-06-06 v1
摘要
互联网已成为仇恨言论(Hate Speech, HS)的热点,威胁社会和谐与个人福祉。虽然自动检测方法在识别显性仇恨言论(ex-HS)方面表现良好,但在识别更隐蔽的形式时,如隐式仇恨言论(im-HS)时却力不从众。为此,我们提出一种新的im-HS检测分类法,定义六种编码策略,称为codetypes。我们呈现两种将codetypes集成到im-HS检测中的方法:1)直接提示大型语言模型(LLMs)根据生成的响应对句子进行分类,以及2)将LLM用作编码器,在编码过程中嵌入codetypes。实验表明,codetypes的使用在中文和英文数据集中均提升了im-HS检测效果,验证了该方法在不同语言中的有效性。
引用
@article{arxiv.2506.04693,
title = {Cracking the Code: Enhancing Implicit Hate Speech Detection through Coding Classification},
author = {Lu Wei and Liangzhi Li and Tong Xiang and Xiao Liu and Noa Garcia},
journal= {arXiv preprint arXiv:2506.04693},
year = {2025}
}
备注
Proceedings of the 5th Workshop on Trustworthy NLP (TrustNLP 2025), 112-126