English

ARPA: Armenian Paraphrase Detection Corpus and Models

Computation and Language 2020-09-29 v1

Abstract

In this work, we employ a semi-automatic method based on back translation to generate a sentential paraphrase corpus for the Armenian language. The initial collection of sentences is translated from Armenian to English and back twice, resulting in pairs of lexically distant but semantically similar sentences. The generated paraphrases are then manually reviewed and annotated. Using the method train and test datasets are created, containing 2360 paraphrases in total. In addition, the datasets are used to train and evaluate BERTbased models for detecting paraphrase in Armenian, achieving results comparable to the state-of-the-art of other languages.

Keywords

Cite

@article{arxiv.2009.12615,
  title  = {ARPA: Armenian Paraphrase Detection Corpus and Models},
  author = {Arthur Malajyan and Karen Avetisyan and Tsolak Ghukasyan},
  journal= {arXiv preprint arXiv:2009.12615},
  year   = {2020}
}

Comments

To be published in the proceedings of Ivannikov Memorial Workshop 2020

R2 v1 2026-06-23T18:48:55.972Z