English

ParaCotta: Synthetic Multilingual Paraphrase Corpora from the Most Diverse Translation Sample Pair

Computation and Language 2022-05-11 v1

Abstract

We release our synthetic parallel paraphrase corpus across 17 languages: Arabic, Catalan, Czech, German, English, Spanish, Estonian, French, Hindi, Indonesian, Italian, Dutch, Romanian, Russian, Swedish, Vietnamese, and Chinese. Our method relies only on monolingual data and a neural machine translation system to generate paraphrases, hence simple to apply. We generate multiple translation samples using beam search and choose the most lexically diverse pair according to their sentence BLEU. We compare our generated corpus with the \texttt{ParaBank2}. According to our evaluation, our synthetic paraphrase pairs are semantically similar and lexically diverse.

Keywords

Cite

@article{arxiv.2205.04651,
  title  = {ParaCotta: Synthetic Multilingual Paraphrase Corpora from the Most Diverse Translation Sample Pair},
  author = {Alham Fikri Aji and Tirana Noor Fatyanosa and Radityo Eko Prasojo and Philip Arthur and Suci Fitriany and Salma Qonitah and Nadhifa Zulfa and Tomi Santoso and Mahendra Data},
  journal= {arXiv preprint arXiv:2205.04651},
  year   = {2022}
}

Comments

10 pages, 3 figures, 6 tables. Accepted at PACLIC 2021. (ACL Anthology link: https://aclanthology.org/2021.paclic-1.56/)