We present a novel corpus of 445 human- and computer-generated documents, comprising about 27,000 clauses, annotated for semantic clause types and coherence relations that allow for nuanced comparison of artificial and natural discourse modes. The corpus covers both formal and informal discourse, and contains documents generated using fine-tuned GPT-2 (Zellers et al., 2019) and GPT-3(Brown et al., 2020). We showcase the usefulness of this corpus for detailed discourse analysis of text generation by providing preliminary evidence that less numerous, shorter and more often incoherent clause relations are associated with lower perceived quality of computer-generated narratives and arguments.
@article{arxiv.2111.05940,
title = {A Novel Corpus of Discourse Structure in Humans and Computers},
author = {Babak Hemmatian and Sheridan Feucht and Rachel Avram and Alexander Wey and Muskaan Garg and Kate Spitalnic and Carsten Eickhoff and Ellie Pavlick and Bjorn Sandstede and Steven Sloman},
journal= {arXiv preprint arXiv:2111.05940},
year = {2021}
}
Comments
In the 2nd Workshop on Computational Approaches to Discourse (CODI) at EMNLP 2021 (extended abstract). 3 pages