English
Related papers

Related papers: Quality of OCR for Degraded Text Images

200 papers

Optical Character Recognition has been a challenging field in the advent of digital computers. It is needed where information is to be readable both to humans and machines. The process of OCR is composed of a set of pre and post processing…

Computer Vision and Pattern Recognition · Computer Science 2018-01-04 Chinmay Chinara , Nishant Nath , Subhajeet Mishra , Sangram Keshari Sahoo , Farida Ashraf Ali

In recent years, text-image joint pre-training techniques have shown promising results in various tasks. However, in Optical Character Recognition (OCR) tasks, aligning text instances with their corresponding text regions in images poses a…

Computer Vision and Pattern Recognition · Computer Science 2024-04-18 Chen Duan , Pei Fu , Shan Guo , Qianyi Jiang , Xiaoming Wei

Precise homography estimation between multiple images is a pre-requisite for many computer vision applications. One application that is particularly relevant in today's digital era is the alignment of scanned or camera-captured document…

Computer Vision and Pattern Recognition · Computer Science 2019-11-15 Kushagra Mahajan , Monika Sharma , Lovekesh Vig

The inherent noise in the observed (e.g., scanned) binary document image degrades the image quality and harms the compression ratio through breaking the pattern repentance and adding entropy to the document images. In this paper, we design…

Computer Vision and Pattern Recognition · Computer Science 2017-04-25 Yandong Guo , Cheng Lu , Jan P. Allebach , Charles A. Bouman

Optical Character Recognition (OCR) technology is widely used to extract text from images of documents, facilitating efficient digitization and data retrieval. However, merely extracting text is insufficient when dealing with complex…

Learned image compression has gained widespread popularity for their efficiency in achieving ultra-low bit-rates. Yet, images containing substantial textual content, particularly screen-content images (SCI), often suffers from text…

Computer Vision and Pattern Recognition · Computer Science 2024-02-14 Chih-Yu Lai , Dung Tran , Kazuhito Koishida

Despite recent advances, standard sequence labeling systems often fail when processing noisy user-generated text or consuming the output of an Optical Character Recognition (OCR) process. In this paper, we improve the noise-aware training…

Computation and Language · Computer Science 2021-05-26 Marcin Namysl , Sven Behnke , Joachim Köhler

Removing noise from scanned pages is a vital step before their submission to the optical character recognition (OCR) system. Most available image denoising methods are supervised where the pairs of noisy/clean pages are required. However,…

Computer Vision and Pattern Recognition · Computer Science 2021-10-11 Mehrdad J Gangeh , Marcin Plata , Hamid Motahari , Nigel P Duffy

The ubiquity of smartphone cameras has led to more and more documents being captured by cameras rather than scanned. Unlike flatbed scanners, photographed documents are often folded and crumpled, resulting in large local variance in text…

Computer Vision and Pattern Recognition · Computer Science 2020-08-06 Amir Markovitz , Inbal Lavi , Or Perel , Shai Mazor , Roee Litman

Documents often exhibit various forms of degradation, which make it hard to be read and substantially deteriorate the performance of an OCR system. In this paper, we propose an effective end-to-end framework named Document Enhancement…

Computer Vision and Pattern Recognition · Computer Science 2020-10-20 Mohamed Ali Souibgui , Yousri Kessentini

We combine three methods which significantly improve the OCR accuracy of OCR models trained on early printed books: (1) The pretraining method utilizes the information stored in already existing models trained on a variety of typesets…

Computer Vision and Pattern Recognition · Computer Science 2018-03-01 Christian Reul , Uwe Springmann , Christoph Wick , Frank Puppe

Historical documents frequently suffer from damage and inconsistencies, including missing or illegible text resulting from issues such as holes, ink problems, and storage damage. These missing portions or gaps are referred to as lacunae. In…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Jaydeep Borkar , David A. Smith

Having a reliable accuracy score is crucial for real world applications of OCR, since such systems are judged by the number of false readings. Lexicon-based OCR systems, which deal with what is essentially a multi-class classification…

Computer Vision and Pattern Recognition · Computer Science 2018-07-17 Noam Mor , Lior Wolf

This paper proposes a combination of a convolutional and a LSTM network to improve the accuracy of OCR on early printed books. While the standard model of line based OCR uses a single LSTM layer, we utilize a CNN- and Pooling-Layer…

Computer Vision and Pattern Recognition · Computer Science 2018-02-28 Christoph Wick , Christian Reul , Frank Puppe

The paper presents an automated software tool for lossy compression of grayscale images. Its structure and facilities are described. The tool allows compressing images by different coders according to a chosen metric from an available set…

Image and Video Processing · Electrical Eng. & Systems 2024-08-29 Sergey Krivenko , Alexander Zemliachenko , Vladimir Lukin , Alexander Zelensky

We propose a simple method for estimating noise level from a single color image. In most image-denoising algorithms, an accurate noise-level estimate results in good denoising performance; however, it is difficult to estimate noise level…

Computer Vision and Pattern Recognition · Computer Science 2019-04-05 Akihiro Nakamura , Michihiro Kobayashi

Noise is always presents in digital images during image acquisition, coding, transmission, and processing steps. Noise is very difficult to remove it from the digital images without the prior knowledge of noise model. That is why, review of…

Computer Vision and Pattern Recognition · Computer Science 2015-05-14 Ajay Kumar Boyat , Brijendra Kumar Joshi

Optical Character Recognition (OCR) is an established task with the objective of identifying the text present in an image. While many off-the-shelf OCR models exist, they are often trained for either scientific (e.g., formulae) or generic…

Computation and Language · Computer Science 2024-03-26 Nan Zhang , Connor Heaton , Sean Timothy Okonsky , Prasenjit Mitra , Hilal Ezgi Toraman

The purpose of this study is to explore the performance of Informed OCR or iOCR. iOCR was developed with a spell correction algorithm to fix errors introduced by conventional OCR for vote tabulation. The results found that the iOCR system…

Emerging Technologies · Computer Science 2022-08-02 Kenneth U. Oyibo , Jean D. Louis , Juan E. Gilbert

In this paper, we investigate the usage of fine-grained font recognition on OCR for books printed from the 15th to the 18th century. We used a newly created dataset for OCR of early printed books for which fonts are labeled with bounding…

Computer Vision and Pattern Recognition · Computer Science 2024-03-27 Mathias Seuret , Janne van der Loop , Nikolaus Weichselbaumer , Martin Mayr , Janina Molnar , Tatjana Hass , Florian Kordon , Anguelos Nicolau , Vincent Christlein