中文

用于幽默分析的众包标注西班牙语语料库

计算与语言 2018-07-20 v4

摘要

计算幽默涉及多项任务,例如幽默识别、幽默生成和幽默评分,这些任务都需要人工整理的数据。在本研究中,我们提出了一个包含 27,000 条西班牙语推文的语料库,这些推文通过众包标注了其幽默价值和趣味性评分,每条推文约有四条标注,由互联网上的 1,300 人标记。该语料库平均分为来自幽默账户和非幽默账户的推文。标注者间一致性 Krippendorff's alpha 值为 0.5710。该数据集可供通用,可作为幽默检测的基础,也是处理主观性的第一步。

关键词

引用

@article{arxiv.1710.00477,
  title  = {A Crowd-Annotated Spanish Corpus for Humor Analysis},
  author = {Santiago Castro and Luis Chiruzzo and Aiala Rosá and Diego Garat and Guillermo Moncecchi},
  journal= {arXiv preprint arXiv:1710.00477},
  year   = {2018}
}

备注

Camera-ready version of the paper submitted to SocialNLP 2018, with a fixed typo