中文

Twitter 数据上无监督文本表示方法的经验性综述

计算与语言 2020-12-08 v1 机器学习

摘要

近年来 NLP 领域取得了前所未有的成就。最值得注意的是,随着基于 Transformer 的大规模预训练语言模型(如 BERT)的出现,文本表示有了显著提升。然而,这些改进是否能迁移到推文这类含噪声的用户生成文本上仍不明确。本文针对噪声 Twitter 数据上的文本聚类任务,对多种知名文本表示技术进行了实验性综述。我们的结果表明,更先进的模型未必在推文上表现最佳,该领域尚需更多探索。

关键词

引用

@article{arxiv.2012.03468,
  title  = {An Empirical Survey of Unsupervised Text Representation Methods on Twitter Data},
  author = {Lili Wang and Chongyang Gao and Jason Wei and Weicheng Ma and Ruibo Liu and Soroush Vosoughi},
  journal= {arXiv preprint arXiv:2012.03468},
  year   = {2020}
}

备注

In proceedings of the 6th Workshop on Noisy User-generated Text (W-NUT) at EMNLP 2020