English

MinWikiSplit: A Sentence Splitting Corpus with Minimal Propositions

Computation and Language 2019-09-27 v1

Abstract

We compiled a new sentence splitting corpus that is composed of 203K pairs of aligned complex source and simplified target sentences. Contrary to previously proposed text simplification corpora, which contain only a small number of split examples, we present a dataset where each input sentence is broken down into a set of minimal propositions, i.e. a sequence of sound, self-contained utterances with each of them presenting a minimal semantic unit that cannot be further decomposed into meaningful propositions. This corpus is useful for developing sentence splitting approaches that learn how to transform sentences with a complex linguistic structure into a fine-grained representation of short sentences that present a simple and more regular structure which is easier to process for downstream applications and thus facilitates and improves their performance.

Keywords

Cite

@article{arxiv.1909.12131,
  title  = {MinWikiSplit: A Sentence Splitting Corpus with Minimal Propositions},
  author = {Christina Niklaus and Andre Freitas and Siegfried Handschuh},
  journal= {arXiv preprint arXiv:1909.12131},
  year   = {2019}
}
R2 v1 2026-06-23T11:26:58.475Z