中文

利用数百万表情符号出现学习跨领域表示以检测情感、情绪与讽刺

机器学习 2019-11-19 v2 机器学习

摘要

NLP 任务常常受限于人工标注数据的稀缺。在社交媒体情感分析及相关任务中,研究人员因此使用二元化表情符号和特定标签作为远程监督的形式。本文表明,通过将远程监督扩展到更多样化的噪声标签集,模型可以学习到更丰富的表示。通过对包含 64 个常见表情符号之一的 12.46 亿条推文数据集进行表情符号预测,我们使用单一预训练模型在情感、情绪和讽刺检测的 8 个基准数据集上取得了 SOTA 性能。我们的分析证实,情感标签的多样性相比以往的远程监督方法带来了性能提升。

关键词

引用

@article{arxiv.1708.00524,
  title  = {Using millions of emoji occurrences to learn any-domain representations for detecting sentiment, emotion and sarcasm},
  author = {Bjarke Felbo and Alan Mislove and Anders Søgaard and Iyad Rahwan and Sune Lehmann},
  journal= {arXiv preprint arXiv:1708.00524},
  year   = {2019}
}

备注

Accepted at EMNLP 2017. Please include EMNLP in any citations. Minor changes from the EMNLP camera-ready version. 9 pages + references and supplementary material