English

Automatic Identification of Document Translations in Large Multilingual Document Collections

Computation and Language 2007-05-23 v1 Information Retrieval

Abstract

Texts and their translations are a rich linguistic resource that can be used to train and test statistics-based Machine Translation systems and many other applications. In this paper, we present a working system that can identify translations and other very similar documents among a large number of candidates, by representing the document contents with a vector of thesaurus terms from a multilingual thesaurus, and by then measuring the semantic similarity between the vectors. Tests on different text types have shown that the system can detect translations with over 96% precision in a large search space of 820 documents or more. The system was tuned to ignore language-specific similarities and to give similar documents in a second language the same similarity score as equivalent documents in the same language. The application can also be used to detect cross-lingual document plagiarism.

Keywords

Cite

@article{arxiv.cs/0609060,
  title  = {Automatic Identification of Document Translations in Large Multilingual Document Collections},
  author = {Bruno Pouliquen and Ralf Steinberger and Camelia Ignat},
  journal= {arXiv preprint arXiv:cs/0609060},
  year   = {2007}
}

Comments

This technology is used daily to link related news items across languages in the multilingual news analysis system NewsExplorer, which is freely accessible at http://press.jrc.it/NewsExplorer . 8 pages

R2 v1 2026-07-22T12:26:39.643Z