English

TEIMMA: The First Content Reuse Annotator for Text, Images, and Math

Information Retrieval 2023-06-14 v2

Abstract

This demo paper presents the first tool to annotate the reuse of text, images, and mathematical formulae in a document pair -- TEIMMA. Annotating content reuse is particularly useful to develop plagiarism detection algorithms. Real-world content reuse is often obfuscated, which makes it challenging to identify such cases. TEIMMA allows entering the obfuscation type to enable novel classifications for confirmed cases of plagiarism. It enables recording different reuse types for text, images, and mathematical formulae in HTML and supports users by visualizing the content reuse in a document pair using similarity detection methods for text and math.

Keywords

Cite

@article{arxiv.2305.13193,
  title  = {TEIMMA: The First Content Reuse Annotator for Text, Images, and Math},
  author = {Ankit Satpute and André Greiner-Petter and Moritz Schubotz and Norman Meuschke and Akiko Aizawa and Olaf Teschke and Bela Gipp},
  journal= {arXiv preprint arXiv:2305.13193},
  year   = {2023}
}
R2 v1 2026-06-28T10:41:40.318Z