中文

用户级聚合对基于词汇的群体级预测的显著益处

社会与信息网络 2018-08-30 v1 计算机与社会

摘要

基于社交媒体文本的现时预测有望对社区级结果提供无侵扰且近实时的预测。这些结果通常关乎人群,但数据在聚合时往往不顾及每个社区 Twitter 人口中的用户。本文描述了一种简单而有效的方法,利用按用户聚合的 Twitter 语言构建社区级模型。在四项不同的美国县级任务(涵盖人口、健康与心理结果)上的结果显示,相较于聚合所有推文的标准做法,预测精度获得巨大且一致的提升(例如中位数收入预测 Pearson r 从 .73 升至 .82,生活满意度预测 r 从 .37 升至 .47)。我们将由 370 亿条推文——其中逾 10 亿条被映射到县——衍生的聚合且匿名化的社区级数据开放供研究使用。

关键词

引用

@article{arxiv.1808.09600,
  title  = {The Remarkable Benefit of User-Level Aggregation for Lexical-based Population-Level Predictions},
  author = {Salvatore Giorgi and Daniel Preotiuc-Pietro and Anneke Buffone and Daniel Rieman and Lyle H. Ungar and H. Andrew Schwartz},
  journal= {arXiv preprint arXiv:1808.09600},
  year   = {2018}
}

备注

To appear in the proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP)