English

Decay No More: A Persistent Twitter Dataset for Learning Social Meaning

Computation and Language 2022-05-10 v2 Artificial Intelligence

Abstract

With the proliferation of social media, many studies resort to social media to construct datasets for developing social meaning understanding systems. For the popular case of Twitter, most researchers distribute tweet IDs without the actual text contents due to the data distribution policy of the platform. One issue is that the posts become increasingly inaccessible over time, which leads to unfair comparisons and a temporal bias in social media research. To alleviate this challenge of data decay, we leverage a paraphrase model to propose a new persistent English Twitter dataset for social meaning (PTSM). PTSM consists of 1717 social meaning datasets in 1010 categories of tasks. We experiment with two SOTA pre-trained language models and show that our PTSM can substitute the actual tweets with paraphrases with marginal performance loss.

Keywords

Cite

@article{arxiv.2204.04611,
  title  = {Decay No More: A Persistent Twitter Dataset for Learning Social Meaning},
  author = {Chiyu Zhang and Muhammad Abdul-Mageed and El Moatez Billah Nagoudi},
  journal= {arXiv preprint arXiv:2204.04611},
  year   = {2022}
}

Comments

1st Workshop on Novel Evaluation Approaches for Text Classification Systems on Social Media (NEATCLasS) colocated at ICWSM 2022. arXiv admin note: text overlap with arXiv:2108.00356

R2 v1 2026-06-24T10:43:30.362Z