中文

预测社交媒体中攻击性帖子的类型与目标

计算与语言 2019-04-17 v2

摘要

随着攻击性内容在社交媒体中变得普遍,已有大量研究致力于识别潜在的攻击性消息。然而,以往关于该主题的工作并未将问题作为一个整体来考虑,而是侧重于检测非常特定类型的攻击性内容,例如仇恨言论、网络欺凌或网络攻击。相比之下,我们在此针对几种不同类型的攻击性内容。具体而言,我们以分层方式建模该任务,识别社交媒体中攻击性消息的类型与目标。为此,我们整理了Offensive Language Identification Dataset (OLID),这是一个新的数据集,其中的推文使用细粒度的三层标注方案标注了攻击性内容,我们已将其公开。我们讨论了OLID与已有的用于仇恨言论识别、攻击性检测及类似任务的数据集之间的主要异同。我们进一步进行了实验,并比较了不同机器学习模型在OLID上的性能。

关键词

引用

@article{arxiv.1902.09666,
  title  = {Predicting the Type and Target of Offensive Posts in Social Media},
  author = {Marcos Zampieri and Shervin Malmasi and Preslav Nakov and Sara Rosenthal and Noura Farra and Ritesh Kumar},
  journal= {arXiv preprint arXiv:1902.09666},
  year   = {2019}
}

备注

Proceedings of the 2019 Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL)