English

The Remarkable Benefit of User-Level Aggregation for Lexical-based Population-Level Predictions

Social and Information Networks 2018-08-30 v1 Computers and Society

Abstract

Nowcasting based on social media text promises to provide unobtrusive and near real-time predictions of community-level outcomes. These outcomes are typically regarding people, but the data is often aggregated without regard to users in the Twitter populations of each community. This paper describes a simple yet effective method for building community-level models using Twitter language aggregated by user. Results on four different U.S. county-level tasks, spanning demographic, health, and psychological outcomes show large and consistent improvements in prediction accuracies (e.g. from Pearson r=.73 to .82 for median income prediction or r=.37 to .47 for life satisfaction prediction) over the standard approach of aggregating all tweets. We make our aggregated and anonymized community-level data, derived from 37 billion tweets -- over 1 billion of which were mapped to counties, available for research.

Keywords

Cite

@article{arxiv.1808.09600,
  title  = {The Remarkable Benefit of User-Level Aggregation for Lexical-based Population-Level Predictions},
  author = {Salvatore Giorgi and Daniel Preotiuc-Pietro and Anneke Buffone and Daniel Rieman and Lyle H. Ungar and H. Andrew Schwartz},
  journal= {arXiv preprint arXiv:1808.09600},
  year   = {2018}
}

Comments

To appear in the proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (EMNLP)