中文
相关论文

相关论文: Cleaning Dirty Books: Post-OCR Processing for Prev…

200 篇论文

Document denoising is considered one of the most challenging tasks in computer vision. There exist millions of documents that are still to be digitized, but problems like document degradation due to natural and man-made factors make this…

计算机视觉与模式识别 · 计算机科学 2023-07-06 Yashowardhan Shinde , Kishore Kulkarni , Sachin Kuberkar

A great deal of historical corpora suffer from errors introduced by the OCR (optical character recognition) methods used in the digitization process. Correcting these errors manually is a time-consuming process and a great part of the…

计算与语言 · 计算机科学 2020-07-23 Mika Hämäläinen , Simon Hengchen

Optical Character Recognition (OCR) technology has revolutionized the digitization of printed text, enabling efficient data extraction and analysis across various domains. Just like Machine Translation systems, OCR systems are prone to…

计算与语言 · 计算机科学 2025-02-25 Harshvivek Kashid , Pushpak Bhattacharyya

Kurdish libraries have many historical publications that were printed back in the early days when printing devices were brought to Kurdistan. Having a good Optical Character Recognition (OCR) to help process these publications and…

计算与语言 · 计算机科学 2024-04-10 Blnd Yaseen , Hossein Hassani

This paper explores the use of a learned classifier for post-OCR text correction. Experiments with the Arabic language show that this approach, which integrates a weighted confusion matrix and a shallow language model, improves the vast…

信息检索 · 计算机科学 2020-06-11 Ido Kissos , Nachum Dershowitz

This article describes the results of a case study that applies Neural Network-based Optical Character Recognition (OCR) to scanned images of books printed between 1487 and 1870 by training the OCR engine OCRopus [@breuel2013high] on the…

计算与语言 · 计算机科学 2017-04-11 U. Springmann , A. Lüdeling

We present a new method to detect duplicates used to merge different bibliographic record corpora with the help of lexical and social information. As we show, a trivial key is not available to delete useless documents. Merging heteregeneous…

数据库 · 计算机科学 2015-04-29 Nicolas Turenne

Optical Character Recognition (OCR) plays a crucial role in digitizing historical and multilingual documents, yet OCR errors - imperfect extraction of text, including character insertion, deletion, and substitution can significantly impact…

计算与语言 · 计算机科学 2025-09-22 Bhawna Piryani , Jamshid Mozafari , Abdelrahman Abdallah , Antoine Doucet , Adam Jatowt

In Document Understanding, the challenge of reconstructing damaged, occluded, or incomplete text remains a critical yet unexplored problem. Subsequent document understanding tasks can benefit from a document reconstruction process. In…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Kunal Purkayastha , Ayan Banerjee , Josep Llados , Umapada Pal

Document comparison typically relies on optical character recognition (OCR) as its core technology. However, OCR requires the selection of appropriate language models for each document and the performance of multilingual or hybrid models…

计算机视觉与模式识别 · 计算机科学 2024-12-06 Doyoung Park , Naresh Reddy Yarram , Sunjin Kim , Minkyu Kim , Seongho Cho , Taehee Lee

This paper investigates the uncertainty of Generative Pre-trained Transformer (GPT) models in extracting mathematical equations from images of varying resolutions and converting them into LaTeX code. We employ concepts of entropy and mutual…

信息论 · 计算机科学 2024-12-10 Alexei Kaltchenko

Many studies on (Offline) Handwritten Text Recognition (HTR) systems have focused on building state-of-the-art models for line recognition on small corpora. However, adding HTR capability to a large scale multilingual OCR system poses new…

计算机视觉与模式识别 · 计算机科学 2019-06-18 R. Reeve Ingle , Yasuhisa Fujii , Thomas Deselaers , Jonathan Baccash , Ashok C. Popat

With the growing adoption of Retrieval-Augmented Generation (RAG) in document processing, robust text recognition has become increasingly critical for knowledge extraction. While OCR (Optical Character Recognition) for English and other…

计算机视觉与模式识别 · 计算机科学 2025-06-30 Ahmed Heakl , Abdullah Sohail , Mukul Ranjan , Rania Hossam , Ghazi Shazan Ahmad , Mohamed El-Geish , Omar Maher , Zhiqiang Shen , Fahad Khan , Salman Khan

Extracting the relevant information out of a large number of documents is a challenging and tedious task. The quality of results generated by the traditionally available full-text search engine and text-based image retrieval systems is not…

信息检索 · 计算机科学 2022-12-05 Riya Gupta , C. V. Jawahar

The lack of large-scale datasets has been a major hindrance to the development of NLP tasks such as spelling correction and grammatical error correction (GEC). As a complementary new resource for these tasks, we present the GitHub Typo…

计算与语言 · 计算机科学 2019-12-02 Masato Hagiwara , Masato Mita

The digitization of historical documents is crucial for preserving the cultural heritage of the society. An important step in this process is converting scanned images to text using Optical Character Recognition (OCR), which can enable…

计算与语言 · 计算机科学 2024-09-04 Angel Beshirov , Milena Dobreva , Dimitar Dimitrov , Momchil Hardalov , Ivan Koychev , Preslav Nakov

The scientific image integrity area presents a challenging research bottleneck, the lack of available datasets to design and evaluate forensic techniques. Its data sensitivity creates a legal hurdle that prevents one to rely on real…

计算机视觉与模式识别 · 计算机科学 2024-09-30 João P. Cardenuto , Anderson Rocha

Contrary to popular belief, Optical Character Recognition (OCR) remains a challenging problem when text occurs in unconstrained environments, like natural scenes, due to geometrical distortions, complex backgrounds, and diverse fonts. In…

计算机视觉与模式识别 · 计算机科学 2019-06-06 Marcin Namysl , Iuliu Konya

Extracting fine-grained OCR text from aged documents in diacritic languages remains challenging due to unexpected artifacts, time-induced degradation, and lack of datasets. While standalone spell correction approaches have been proposed,…

计算与语言 · 计算机科学 2025-02-28 Thao Do , Dinh Phu Tran , An Vo , Daeyoung Kim

In this paper, we investigate the usage of fine-grained font recognition on OCR for books printed from the 15th to the 18th century. We used a newly created dataset for OCR of early printed books for which fonts are labeled with bounding…