English

FakeCovid -- A Multilingual Cross-domain Fact Check News Dataset for COVID-19

Computers and Society 2020-06-23 v1 Social and Information Networks

Abstract

In this paper, we present a first multilingual cross-domain dataset of 5182 fact-checked news articles for COVID-19, collected from 04/01/2020 to 15/05/2020. We have collected the fact-checked articles from 92 different fact-checking websites after obtaining references from Poynter and Snopes. We have manually annotated articles into 11 different categories of the fact-checked news according to their content. The dataset is in 40 languages from 105 countries. We have built a classifier to detect fake news and present results for the automatic fake news detection and its class. Our model achieves an F1 score of 0.76 to detect the false class and other fact check articles. The FakeCovid dataset is available at Github.

Keywords

Cite

@article{arxiv.2006.11343,
  title  = {FakeCovid -- A Multilingual Cross-domain Fact Check News Dataset for COVID-19},
  author = {Gautam Kishore Shahi and Durgesh Nandini},
  journal= {arXiv preprint arXiv:2006.11343},
  year   = {2020}
}

Comments

CySoc 2020 International Workshop on Cyber Social Threats, ICWSM 2020