Deep Investigation of Cross-Language Plagiarism Detection Methods
Computation and Language
2017-05-25 v1
Abstract
This paper is a deep investigation of cross-language plagiarism detection methods on a new recently introduced open dataset, which contains parallel and comparable collections of documents with multiple characteristics (different genres, languages and sizes of texts). We investigate cross-language plagiarism detection methods for 6 language pairs on 2 granularities of text units in order to draw robust conclusions on the best methods while deeply analyzing correlations across document styles and languages.
Keywords
Cite
@article{arxiv.1705.08828,
title = {Deep Investigation of Cross-Language Plagiarism Detection Methods},
author = {Jeremy Ferrero and Laurent Besacier and Didier Schwab and Frederic Agnes},
journal= {arXiv preprint arXiv:1705.08828},
year = {2017}
}
Comments
Accepted to BUCC (10th Workshop on Building and Using Comparable Corpora) colocated with ACL 2017