Twitter 数据上无监督文本表示方法的经验性综述
计算与语言
2020-12-08 v1 机器学习
摘要
近年来 NLP 领域取得了前所未有的成就。最值得注意的是,随着基于 Transformer 的大规模预训练语言模型(如 BERT)的出现,文本表示有了显著提升。然而,这些改进是否能迁移到推文这类含噪声的用户生成文本上仍不明确。本文针对噪声 Twitter 数据上的文本聚类任务,对多种知名文本表示技术进行了实验性综述。我们的结果表明,更先进的模型未必在推文上表现最佳,该领域尚需更多探索。
引用
@article{arxiv.2012.03468,
title = {An Empirical Survey of Unsupervised Text Representation Methods on Twitter Data},
author = {Lili Wang and Chongyang Gao and Jason Wei and Weicheng Ma and Ruibo Liu and Soroush Vosoughi},
journal= {arXiv preprint arXiv:2012.03468},
year = {2020}
}
备注
In proceedings of the 6th Workshop on Noisy User-generated Text (W-NUT) at EMNLP 2020