中文

从有机文本中甄别机器人文本:一种检测Twitter自动化的自然语言方法

计算与语言 2016-06-15 v6

摘要

Twitter作为流行的社交媒体平台,已演变为丰富的语言数据来源,充满观点、情感和讨论。由于Twitter日益流行,其潜在的社会影响力导致了多样化自动程序社区(通常称为bots)的兴起。这些非有机和半有机Twitter实体范围从良性(如天气更新bots、求职警报bots)到恶意(如垃圾消息、广告或极端观点)。现有检测算法通常利用元数据(推文间隔、粉丝数等)识别机器人账户。本文提出一种强大的分类方案,专门使用来自有机用户的自然语言文本,为识别发布自动消息的账户提供标准。由于分类器仅基于文本运行,它具有灵活性,可应用于Twitter圈之外的任何文本数据。

关键词

引用

@article{arxiv.1505.04342,
  title  = {Sifting Robotic from Organic Text: A Natural Language Approach for Detecting Automation on Twitter},
  author = {Eric M. Clark and Jake Ryland Williams and Chris A. Jones and Richard A. Galbraith and Christopher M. Danforth and Peter Sheridan Dodds},
  journal= {arXiv preprint arXiv:1505.04342},
  year   = {2016}
}