English

A Tweet-based Dataset for Company-Level Stock Return Prediction

Computation and Language 2020-06-18 v1 Social and Information Networks Statistical Finance

Abstract

Public opinion influences events, especially related to stock market movement, in which a subtle hint can influence the local outcome of the market. In this paper, we present a dataset that allows for company-level analysis of tweet based impact on one-, two-, three-, and seven-day stock returns. Our dataset consists of 862, 231 labelled instances from twitter in English, we also release a cleaned subset of 85, 176 labelled instances to the community. We also provide baselines using standard machine learning algorithms and a multi-view learning based approach that makes use of different types of features. Our dataset, scripts and models are publicly available at: https://github.com/ImperialNLP/stockreturnpred.

Keywords

Cite

@article{arxiv.2006.09723,
  title  = {A Tweet-based Dataset for Company-Level Stock Return Prediction},
  author = {Karolina Sowinska and Pranava Madhyastha},
  journal= {arXiv preprint arXiv:2006.09723},
  year   = {2020}
}

Comments

Dataset available here: https://github.com/ImperialNLP/stockreturnpred

R2 v1 2026-06-23T16:23:52.793Z