English

A Simple and Effective Method of Cross-Lingual Plagiarism Detection

Computation and Language 2023-04-06 v2 Information Retrieval

Abstract

We present a simple cross-lingual plagiarism detection method applicable to a large number of languages. The presented approach leverages open multilingual thesauri for candidate retrieval task and pre-trained multilingual BERT-based language models for detailed analysis. The method does not rely on machine translation and word sense disambiguation when in use, and therefore is suitable for a large number of languages, including under-resourced languages. The effectiveness of the proposed approach is demonstrated for several existing and new benchmarks, achieving state-of-the-art results for French, Russian, and Armenian languages.

Keywords

Cite

@article{arxiv.2304.01352,
  title  = {A Simple and Effective Method of Cross-Lingual Plagiarism Detection},
  author = {Karen Avetisyan and Arthur Malajyan and Tsolak Ghukasyan and Arutyun Avetisyan},
  journal= {arXiv preprint arXiv:2304.01352},
  year   = {2023}
}