English

RusTitW: Russian Language Text Dataset for Visual Text in-the-Wild Recognition

Computer Vision and Pattern Recognition 2023-03-30 v1

Abstract

Information surrounds people in modern life. Text is a very efficient type of information that people use for communication for centuries. However, automated text-in-the-wild recognition remains a challenging problem. The major limitation for a DL system is the lack of training data. For the competitive performance, training set must contain many samples that replicate the real-world cases. While there are many high-quality datasets for English text recognition; there are no available datasets for Russian language. In this paper, we present a large-scale human-labeled dataset for Russian text recognition in-the-wild. We also publish a synthetic dataset and code to reproduce the generation process

Keywords

Cite

@article{arxiv.2303.16531,
  title  = {RusTitW: Russian Language Text Dataset for Visual Text in-the-Wild Recognition},
  author = {Igor Markov and Sergey Nesteruk and Andrey Kuznetsov and Denis Dimitrov},
  journal= {arXiv preprint arXiv:2303.16531},
  year   = {2023}
}

Comments

5 pages, 6 figures, 2 tables

R2 v1 2026-06-28T09:39:27.502Z