English

Learning to Simplify with Data Hopelessly Out of Alignment

Computation and Language 2022-04-05 v1

Abstract

We consider whether it is possible to do text simplification without relying on a "parallel" corpus, one that is made up of sentence-by-sentence alignments of complex and ground truth simple sentences. To this end, we introduce a number of concepts, some new and some not, including what we call Conjoined Twin Networks, Flip-Flop Auto-Encoders (FFA) and Adversarial Networks (GAN). A comparison is made between Jensen-Shannon (JS-GAN) and Wasserstein GAN, to see how they impact performance, with stronger results for the former. An experiment we conducted with a large dataset derived from Wikipedia found the solid superiority of Twin Networks equipped with FFA and JS-GAN, over the current best performing system. Furthermore, we discuss where we stand in a relation to fully supervised methods in the past literature, and highlight with examples qualitative differences that exist among simplified sentences generated by supervision-free systems.

Keywords

Cite

@article{arxiv.2204.00741,
  title  = {Learning to Simplify with Data Hopelessly Out of Alignment},
  author = {Tadashi Nomoto},
  journal= {arXiv preprint arXiv:2204.00741},
  year   = {2022}
}
R2 v1 2026-06-24T10:35:19.511Z