English

RETSim: Resilient and Efficient Text Similarity

Computation and Language 2023-11-30 v1

Abstract

This paper introduces RETSim (Resilient and Efficient Text Similarity), a lightweight, multilingual deep learning model trained to produce robust metric embeddings for near-duplicate text retrieval, clustering, and dataset deduplication tasks. We demonstrate that RETSim is significantly more robust and accurate than MinHash and neural text embeddings, achieving new state-of-the-art performance on dataset deduplication, adversarial text retrieval benchmarks, and spam clustering tasks. We also introduce the W4NT3D benchmark (Wiki-40B 4dversarial Near-T3xt Dataset) for evaluating multilingual, near-duplicate text retrieval capabilities under adversarial settings. RETSim and the W4NT3D benchmark are open-sourced under the MIT License at https://github.com/google/unisim.

Keywords

Cite

@article{arxiv.2311.17264,
  title  = {RETSim: Resilient and Efficient Text Similarity},
  author = {Marina Zhang and Owen Vallis and Aysegul Bumin and Tanay Vakharia and Elie Bursztein},
  journal= {arXiv preprint arXiv:2311.17264},
  year   = {2023}
}
R2 v1 2026-06-28T13:34:50.061Z