English

A Benchmark Dataset of Check-worthy Factual Claims

Computation and Language 2020-05-01 v1

Abstract

In this paper we present the ClaimBuster dataset of 23,533 statements extracted from all U.S. general election presidential debates and annotated by human coders. The ClaimBuster dataset can be leveraged in building computational methods to identify claims that are worth fact-checking from the myriad of sources of digital or traditional media. The ClaimBuster dataset is publicly available to the research community, and it can be found at http://doi.org/10.5281/zenodo.3609356.

Keywords

Cite

@article{arxiv.2004.14425,
  title  = {A Benchmark Dataset of Check-worthy Factual Claims},
  author = {Fatma Arslan and Naeemul Hassan and Chengkai Li and Mark Tremayne},
  journal= {arXiv preprint arXiv:2004.14425},
  year   = {2020}
}

Comments

Accepted to ICWSM 2020