English

COVID-Twitter-BERT: A Natural Language Processing Model to Analyse COVID-19 Content on Twitter

Computation and Language 2020-05-18 v1 Machine Learning Social and Information Networks

Abstract

In this work, we release COVID-Twitter-BERT (CT-BERT), a transformer-based model, pretrained on a large corpus of Twitter messages on the topic of COVID-19. Our model shows a 10-30% marginal improvement compared to its base model, BERT-Large, on five different classification datasets. The largest improvements are on the target domain. Pretrained transformer models, such as CT-BERT, are trained on a specific target domain and can be used for a wide variety of natural language processing tasks, including classification, question-answering and chatbots. CT-BERT is optimised to be used on COVID-19 content, in particular social media posts from Twitter.

Keywords

Cite

@article{arxiv.2005.07503,
  title  = {COVID-Twitter-BERT: A Natural Language Processing Model to Analyse COVID-19 Content on Twitter},
  author = {Martin Müller and Marcel Salathé and Per E Kummervold},
  journal= {arXiv preprint arXiv:2005.07503},
  year   = {2020}
}
R2 v1 2026-06-23T15:34:17.675Z