English

A German Corpus for Text Similarity Detection Tasks

Information Retrieval 2017-03-14 v1 Computation and Language

Abstract

Text similarity detection aims at measuring the degree of similarity between a pair of texts. Corpora available for text similarity detection are designed to evaluate the algorithms to assess the paraphrase level among documents. In this paper we present a textual German corpus for similarity detection. The purpose of this corpus is to automatically assess the similarity between a pair of texts and to evaluate different similarity measures, both for whole documents or for individual sentences. Therefore we have calculated several simple measures on our corpus based on a library of similarity functions.

Keywords

Cite

@article{arxiv.1703.03923,
  title  = {A German Corpus for Text Similarity Detection Tasks},
  author = {Juan-Manuel Torres-Moreno and Gerardo Sierra and Peter Peinl},
  journal= {arXiv preprint arXiv:1703.03923},
  year   = {2017}
}

Comments

1 figure; 13 pages

R2 v1 2026-06-22T18:42:54.956Z