中文
相关论文

相关论文: Upcycle Your OCR: Reusing OCRs for Post-OCR Text C…

200 篇论文

In this paper, we propose a data augmentation framework for Optical Character Recognition (OCR). The proposed framework is able to synthesize new viewing angles and illumination scenarios, effectively enriching any available OCR dataset.…

计算机视觉与模式识别 · 计算机科学 2022-09-30 Andreas Spruck , Maximiliane Hawesch , Anatol Maier , Christian Riess , Jürgen Seiler , André Kaup

This research paper delves into the development of an Optical Character Recognition (OCR) system for the recognition of Ashokan Brahmi characters using Convolutional Neural Networks. It utilizes a comprehensive dataset of character images…

计算机视觉与模式识别 · 计算机科学 2025-01-07 Yash Agrawal , Srinidhi Balasubramanian , Rahul Meena , Rohail Alam , Himanshu Malviya , Rohini P

Long-term OCR services aim to provide high-quality output to their users at competitive costs. It is essential to upgrade the models because of the complex data loaded by the users. The service providers encourage the users who provide data…

计算机视觉与模式识别 · 计算机科学 2022-12-20 Ajoy Mondal , Rohit saluja , C. V. Jawahar

This paper presents our methodology and findings from three tasks across Optical Character Recognition (OCR) and Document Layout Analysis using advanced deep learning techniques. First, for the historical Hebrew fragments of the Dead Sea…

计算机视觉与模式识别 · 计算机科学 2025-08-15 Hylke Westerdijk , Ben Blankenborg , Khondoker Ittehadul Islam

We present an end-to-end trainable approach for Optical Character Recognition (OCR) on printed documents. Specifically, we propose a model that predicts a) a two-dimensional character grid (\emph{chargrid}) representation of a document…

计算机视觉与模式识别 · 计算机科学 2020-02-28 Christian Reisswig , Anoop R Katti , Marco Spinaci , Johannes Höhne

Optical character recognition (OCR) is a process of converting analogue documents into digital using document images. Currently, many commercial and non-commercial OCR systems exist for both handwritten and printed copies for different…

计算机视觉与模式识别 · 计算机科学 2021-05-11 Farisa Benta Safir , Abu Quwsar Ohi , M. F. Mridha , Muhammad Mostafa Monowar , Md. Abdul Hamid

This article describes the results of a case study that applies Neural Network-based Optical Character Recognition (OCR) to scanned images of books printed between 1487 and 1870 by training the OCR engine OCRopus [@breuel2013high] on the…

计算与语言 · 计算机科学 2017-04-11 U. Springmann , A. Lüdeling

This paper explores the use of a learned classifier for post-OCR text correction. Experiments with the Arabic language show that this approach, which integrates a weighted confusion matrix and a shallow language model, improves the vast…

信息检索 · 计算机科学 2020-06-11 Ido Kissos , Nachum Dershowitz

Diacritic characters can be considered as a unique set of characters providing us with adequate and significant clue in identifying a given language with considerably high accuracy. Diacritics, though associated with phonetics often serve…

计算机视觉与模式识别 · 计算机科学 2021-12-07 Shubham Vatsal , Nikhil Arora , Gopi Ramena , Sukumar Moharana , Dhruval Jain , Naresh Purre , Rachit S Munjal

Optical character recognition (OCR) methods have been applied to diverse tasks, e.g., street view text recognition and document analysis. Recently, zero-shot OCR has piqued the interest of the research community because it considers a…

计算机视觉与模式识别 · 计算机科学 2023-08-02 Xiaolei Diao , Daqian Shi , Jian Li , Lida Shi , Mingzhe Yue , Ruihua Qi , Chuntao Li , Hao Xu

Recognition of document images have important applications in restoring old and classical texts. The problem involves quality improvement before passing it to a properly trained OCR to get accurate recognition of the text. The image…

计算机视觉与模式识别 · 计算机科学 2017-02-01 Ram Krishna Pandey , A G Ramakrishnan

A method is presented that significantly reduces the character error rates for OCR text obtained from OCRopus models trained on early printed books when only small amounts of diplomatic transcriptions are available. This is achieved by…

计算机视觉与模式识别 · 计算机科学 2017-12-22 Christian Reul , Christoph Wick , Uwe Springmann , Frank Puppe

Designing Optical Character Recognition (OCR) systems for India requires balancing linguistic diversity, document heterogeneity, and deployment constraints. In this paper, we study two training strategies for building multilingual OCR…

计算机视觉与模式识别 · 计算机科学 2026-02-19 Ali Faraz , Raja Kolla , Ashish Kulkarni , Shubham Agarwal

In this paper we introduce a method that significantly reduces the character error rates for OCR text obtained from OCRopus models trained on early printed books. The method uses a combination of cross fold training and confidence based…

计算机视觉与模式识别 · 计算机科学 2018-07-25 Christian Reul , Uwe Springmann , Christoph Wick , Frank Puppe

Extracting fine-grained OCR text from aged documents in diacritic languages remains challenging due to unexpected artifacts, time-induced degradation, and lack of datasets. While standalone spell correction approaches have been proposed,…

计算与语言 · 计算机科学 2025-02-28 Thao Do , Dinh Phu Tran , An Vo , Daeyoung Kim

The digitization of scanned forms and documents is changing the data sources that enterprises manage. To integrate these new data sources with enterprise data, the current state-of-the-art approach is to convert the images to ASCII text…

数据库 · 计算机科学 2012-01-09 Arun Kumar , Christopher Ré

Over the past few decades, large archives of paper-based documents such as books and newspapers have been digitized using Optical Character Recognition. This technology is error-prone, especially for historical documents. To correct OCR…

计算与语言 · 计算机科学 2023-08-01 Omri Suissa , Avshalom Elmalech , Maayan Zhitomirsky-Geffet

Historical corpora are known to contain errors introduced by OCR (optical character recognition) methods used in the digitization process, often said to be degrading the performance of NLP systems. Correcting these errors manually is a…

计算与语言 · 计算机科学 2020-11-20 Quan Duong , Mika Hämäläinen , Simon Hengchen

Optical Character Recognition (OCR) has been a topic of interest for many years. It is defined as the process of digitizing a document image into its constituent characters. Despite decades of intense research, developing OCR with…

计算机视觉与模式识别 · 计算机科学 2017-10-17 Noman Islam , Zeeshan Islam , Nazia Noor

Text image super-resolution is a challenging yet open research problem in the computer vision community. In particular, low-resolution images hamper the performance of typical optical character recognition (OCR) systems. In this article, we…

计算机视觉与模式识别 · 计算机科学 2015-06-09 Chao Dong , Ximei Zhu , Yubin Deng , Chen Change Loy , Yu Qiao