中文

一种简洁且语言无关却极强的仇恨言论与攻击性内容识别基线系统

计算与语言 2022-02-08 v1

摘要

针对推文中仇恨言论与攻击性内容的自动识别,SATLab 团队提出了一种仅以字符 n-gram 为特征、基于经典监督算法的系统,因而完全语言无关。在对其特征加权和分类器参数进行优化后,该系统在多语言 HASOC 2021 挑战赛中,于英语(易于开发依赖大量外部语言资源的深度学习方法)上达到中等性能水平,但在资源较少的印地语和马拉地语上表现远好。当在这两种语言的三项任务上取平均性能时,它甚至位列第一,优于许多深度学习方法。这些性能表明,它是一个有趣的参考水平,可用于评估使用更复杂方法(如深度学习或利用互补资源)所带来的收益。

关键词

引用

@article{arxiv.2202.02511,
  title  = {A simple language-agnostic yet very strong baseline system for hate speech and offensive content identification},
  author = {Yves Bestgen},
  journal= {arXiv preprint arXiv:2202.02511},
  year   = {2022}
}

备注

A slightly modified version of the paper: "A simple language-agnostic yet strong baseline system for hate speech and offensive content identification. In Working Notes of FIRE 2021 - Forum for Information Retrieval Evaluation (10 p.). ceur-ws.org