English

OTR: Synthesizing Overlay Text Dataset for Text Removal

Computer Vision and Pattern Recognition 2025-10-06 v1

Abstract

Text removal is a crucial task in computer vision with applications such as privacy preservation, image editing, and media reuse. While existing research has primarily focused on scene text removal in natural images, limitations in current datasets hinder out-of-domain generalization or accurate evaluation. In particular, widely used benchmarks such as SCUT-EnsText suffer from ground truth artifacts due to manual editing, overly simplistic text backgrounds, and evaluation metrics that do not capture the quality of generated results. To address these issues, we introduce an approach to synthesizing a text removal benchmark applicable to domains other than scene texts. Our dataset features text rendered on complex backgrounds using object-aware placement and vision-language model-generated content, ensuring clean ground truth and challenging text removal scenarios. The dataset is available at https://huggingface.co/datasets/cyberagent/OTR .

Keywords

Cite

@article{arxiv.2510.02787,
  title  = {OTR: Synthesizing Overlay Text Dataset for Text Removal},
  author = {Jan Zdenek and Wataru Shimoda and Kota Yamaguchi},
  journal= {arXiv preprint arXiv:2510.02787},
  year   = {2025}
}

Comments

This is the author's version of the work. It is posted here for your personal use. Not for redistribution. The definitive Version of Record was published in Proceedings of the 33rd ACM International Conference on Multimedia (MM '25), October 27-31, 2025, Dublin, Ireland, https://doi.org/10.1145/3746027.3758297

R2 v1 2026-07-01T06:14:51.640Z