OTR:用于文本移除的叠加文本数据集合成
计算机视觉与模式识别
2025-10-06 v1
摘要
文本移除是计算机视觉中的关键任务,适用于隐私保护、图像编辑和媒体再利用等应用。虽然现有研究主要聚焦于自然图像中的场景文本移除,但当前数据集在跨域泛化或准确评估方面存在局限。特别是,诸如 SCUT-EnsText 等广泛使用的基准数据集因人工编辑而产生真实感 artifact、文本背景过于简单,以及不能捕捉生成结果质量的评估指标等问题。为此,我们提出一种用于合成适用于除场景文本之外领域的文本移除基准的方法。该数据集使用对象感知放置和视觉-语言模型生成的内容,在复杂背景上渲染文本,确保干净的真实感和具有挑战性的文本移除情形。数据集可在 https://huggingface.co/datasets/cyberagent/OTR 获取。
引用
@article{arxiv.2510.02787,
title = {OTR: Synthesizing Overlay Text Dataset for Text Removal},
author = {Jan Zdenek and Wataru Shimoda and Kota Yamaguchi},
journal= {arXiv preprint arXiv:2510.02787},
year = {2025}
}
备注
This is the author's version of the work. It is posted here for your personal use. Not for redistribution. The definitive Version of Record was published in Proceedings of the 33rd ACM International Conference on Multimedia (MM '25), October 27-31, 2025, Dublin, Ireland, https://doi.org/10.1145/3746027.3758297