English

Deep Investigation of Cross-Language Plagiarism Detection Methods

Computation and Language 2017-05-25 v1

Abstract

This paper is a deep investigation of cross-language plagiarism detection methods on a new recently introduced open dataset, which contains parallel and comparable collections of documents with multiple characteristics (different genres, languages and sizes of texts). We investigate cross-language plagiarism detection methods for 6 language pairs on 2 granularities of text units in order to draw robust conclusions on the best methods while deeply analyzing correlations across document styles and languages.

Keywords

Cite

@article{arxiv.1705.08828,
  title  = {Deep Investigation of Cross-Language Plagiarism Detection Methods},
  author = {Jeremy Ferrero and Laurent Besacier and Didier Schwab and Frederic Agnes},
  journal= {arXiv preprint arXiv:1705.08828},
  year   = {2017}
}

Comments

Accepted to BUCC (10th Workshop on Building and Using Comparable Corpora) colocated with ACL 2017