中文
相关论文

相关论文: synthocr-gen: A synthetic ocr dataset generator fo…

200 篇论文

Optical Character Recognition (OCR) is essential in applications such as document processing, license plate recognition, and intelligent surveillance. However, existing OCR models often underperform in real-world scenarios due to irregular…

计算机视觉与模式识别 · 计算机科学 2025-04-09 Inho Jake Park , Jaehoon Jay Jeong , Ho-Sang Jo

This paper introduces PreP-OCR, a two-stage pipeline that combines document image restoration with semantic-aware post-OCR correction to enhance both visual clarity and textual consistency, thereby improving text extraction from degraded…

计算与语言 · 计算机科学 2025-11-19 Shuhao Guan , Moule Lin , Cheng Xu , Xinyi Liu , Jinman Zhao , Jiexin Fan , Qi Xu , Derek Greene

Document semantic segmentation is a promising avenue that can facilitate document analysis tasks, including optical character recognition (OCR), form classification, and document editing. Although several synthetic datasets have been…

计算机视觉与模式识别 · 计算机科学 2024-05-01 Taylor Archibald , Tony Martinez

Indonesia is one of the most diverse countries linguistically. However, despite this linguistic diversity, Indonesian languages remain underrepresented in Natural Language Processing (NLP) research and technologies. In the past two years,…

This paper contributes a new large-scale dataset for weakly supervised cross-media retrieval, named Twitter100k. Current datasets, such as Wikipedia, NUS Wide and Flickr30k, have two major limitations. First, these datasets are lacking in…

计算机视觉与模式识别 · 计算机科学 2017-03-21 Yuting Hu , Liang Zheng , Yi Yang , Yongfeng Huang

Synthetic data is increasingly critical for contact centers, where privacy constraints and data scarcity limit the availability of real conversations. However, generating synthetic dialogues that are realistic and useful for downstream…

计算与语言 · 计算机科学 2026-02-17 Rishikesh Devanathan , Varun Nathan , Ayush Kumar

The primary obstacle to developing technologies for low-resource languages is the lack of usable data. In this paper, we report the adoption and deployment of 4 technology-driven methods of data collection for Gondi, a low-resource…

We propose RareGraph-Synth, a knowledge-guided, continuous-time diffusion framework that generates realistic yet privacy-preserving synthetic electronic-health-record (EHR) trajectories for ultra-rare diseases. RareGraph-Synth unifies five…

机器学习 · 计算机科学 2025-10-09 Khartik Uppalapati , Shakeel Abdulkareem , Bora Yimenicioglu

Developing culturally grounded multilingual AI systems remains challenging, particularly for low-resource languages. While synthetic data offers promise, its effectiveness in multilingual and multicultural contexts is underexplored. We…

Deciphering oracle bone scripts plays an important role in Chinese archaeology and philology. However, a significant challenge remains due to the scarcity of oracle character images. To overcome this issue, we propose Diff-Oracle, a novel…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Jing Li , Qiu-Feng Wang , Siyuan Wang , Rui Zhang , Kaizhu Huang , Erik Cambria

The Tajik language, written in Cyrillic script, remains severely under-resourced in terms of publicly available natural language processing (NLP) toolkits, hindering both linguistic research and applied development. This paper introduces…

计算与语言 · 计算机科学 2026-05-29 Mullosharaf K. Arabov

This research presents a few-shot voice cloning system for Nepali speakers, designed to synthesize speech in a specific speaker's voice from Devanagari text using minimal data. Voice cloning in Nepali remains largely unexplored due to its…

声音 · 计算机科学 2026-01-27 Aayush M. Shrestha , Aditya Bajracharya , Projan Shakya , Dinesh B. Kshatri

Despite remarkable advances in natural language processing, developing effective systems for low-resource languages remains a formidable challenge, with performances typically lagging far behind high-resource counterparts due to data…

人工智能 · 计算机科学 2026-02-06 Subhadip Maji , Arnab Bhattacharya

Generative diffusion models have emerged as powerful tools to synthetically produce training data, offering potential solutions to data scarcity and reducing labelling costs for downstream supervised deep learning applications. However,…

计算机视觉与模式识别 · 计算机科学 2025-04-09 Nicolo Resmini , Eugenio Lomurno , Cristian Sbrolli , Matteo Matteucci

Diffusion-based scene text synthesis has progressed rapidly, yet existing methods commonly rely on additional visual conditioning modules and require large-scale annotated data to support multilingual generation. In this work, we revisit…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Yu Xie , Jielei Zhang , Pengyu Chen , Weihang Wang , Longwen Gao , Peiyi Li , Qian Qiao , Zhouhui Lian

Handwritten Text Recognition (HTR) is a well-established research area. In contrast, Handwritten Text Generation (HTG) is an emerging field with significant potential. This task is challenging due to the variation in individual handwriting…

计算机视觉与模式识别 · 计算机科学 2025-12-29 Md. Rakibul Islam , Md. Kamrozzaman Bhuiyan , Safwan Muntasir , Arifur Rahman Jawad , Most. Sharmin Sultana Samu

In this paper we evaluate Optical Character Recognition (OCR) of 19th century Fraktur scripts without book-specific training using mixed models, i.e. models trained to recognize a variety of fonts and typesets from previously unseen…

计算机视觉与模式识别 · 计算机科学 2018-10-09 Christian Reul , Uwe Springmann , Christoph Wick , Frank Puppe

This research digitizes and analyzes the Leidse hoogleraren en lectoren 1575-1815 books written between 1983 and 1985, which contain biographic data about professors and curators of Leiden University. It addresses the central question: how…

计算与语言 · 计算机科学 2026-01-01 Zahra Abedi , Richard M. K. van Dijk , Gijs Wijnholds , Tessa Verhoef

Recognition of text on word or line images, without the need for sub-word segmentation has become the mainstream of research and development of text recognition for Indian languages. Modelling unsegmented sequences using Connectionist…

计算机视觉与模式识别 · 计算机科学 2024-12-19 Minesh Mathew , Ajoy Mondal , CV Jawahar

In the present work, we have used Tesseract 2.01 open source Optical Character Recognition (OCR) Engine under Apache License 2.0 for recognition of handwriting samples of lower case Roman script. Handwritten isolated and free-flow text…

计算机视觉与模式识别 · 计算机科学 2010-03-31 Sandip Rakshit , Subhadip Basu
‹ 上一页 1 8 9 10 下一页 ›