English

Assessing The Factual Accuracy of Generated Text

Computation and Language 2021-05-27 v2

Abstract

We propose a model-based metric to estimate the factual accuracy of generated text that is complementary to typical scoring schemes like ROUGE (Recall-Oriented Understudy for Gisting Evaluation) and BLEU (Bilingual Evaluation Understudy). We introduce and release a new large-scale dataset based on Wikipedia and Wikidata to train relation classifiers and end-to-end fact extraction models. The end-to-end models are shown to be able to extract complete sets of facts from datasets with full pages of text. We then analyse multiple models that estimate factual accuracy on a Wikipedia text summarization task, and show their efficacy compared to ROUGE and other model-free variants by conducting a human evaluation study.

Keywords

Cite

@article{arxiv.1905.13322,
  title  = {Assessing The Factual Accuracy of Generated Text},
  author = {Ben Goodrich and Vinay Rao and Mohammad Saleh and Peter J Liu},
  journal= {arXiv preprint arXiv:1905.13322},
  year   = {2021}
}
R2 v1 2026-06-23T09:34:10.104Z