English

Improving the Diversity of Unsupervised Paraphrasing with Embedding Outputs

Computation and Language 2021-10-27 v1

Abstract

We present a novel technique for zero-shot paraphrase generation. The key contribution is an end-to-end multilingual paraphrasing model that is trained using translated parallel corpora to generate paraphrases into "meaning spaces" -- replacing the final softmax layer with word embeddings. This architectural modification, plus a training procedure that incorporates an autoencoding objective, enables effective parameter sharing across languages for more fluent monolingual rewriting, and facilitates fluency and diversity in generation. Our continuous-output paraphrase generation models outperform zero-shot paraphrasing baselines when evaluated on two languages using a battery of computational metrics as well as in human assessment.

Keywords

Cite

@article{arxiv.2110.13231,
  title  = {Improving the Diversity of Unsupervised Paraphrasing with Embedding Outputs},
  author = {Monisha Jegadeesan and Sachin Kumar and John Wieting and Yulia Tsvetkov},
  journal= {arXiv preprint arXiv:2110.13231},
  year   = {2021}
}
R2 v1 2026-06-24T07:10:40.366Z