中文
相关论文

相关论文: Enriching Historical Records: An OCR and AI-Driven…

200 篇论文

Retrieving accurate details from documents is a crucial task, especially when handling a combination of scanned images and native digital formats. This document presents a combined framework for text extraction that merges Optical Character…

计算机视觉与模式识别 · 计算机科学 2025-06-16 Rasha Sinha , Rekha B S

Academic documents are packed with texts, equations, tables, and figures, requiring comprehensive understanding for accurate Optical Character Recognition (OCR). While end-to-end OCR methods offer improved accuracy over layout-based…

计算机视觉与模式识别 · 计算机科学 2024-03-05 Yu Sun , Dongzhan Zhou , Chen Lin , Conghui He , Wanli Ouyang , Han-Sen Zhong

Optical Character Recognition (OCR) plays a crucial role in digitizing historical and multilingual documents, yet OCR errors - imperfect extraction of text, including character insertion, deletion, and substitution can significantly impact…

计算与语言 · 计算机科学 2025-09-22 Bhawna Piryani , Jamshid Mozafari , Abdelrahman Abdallah , Antoine Doucet , Adam Jatowt

The digitisation of historical print media archives is crucial for increasing accessibility to contemporary records. However, the process of Optical Character Recognition (OCR) used to convert physical records to digital text is prone to…

计算与语言 · 计算机科学 2025-01-23 Jonathan Bourne

This paper introduces PreP-OCR, a two-stage pipeline that combines document image restoration with semantic-aware post-OCR correction to enhance both visual clarity and textual consistency, thereby improving text extraction from degraded…

计算与语言 · 计算机科学 2025-11-19 Shuhao Guan , Moule Lin , Cheng Xu , Xinyi Liu , Jinman Zhao , Jiexin Fan , Qi Xu , Derek Greene

Digitization of historical documents is a challenging task in many digital humanities projects. A popular approach for digitization is to scan the documents into images, and then convert images into text using Optical Character Recognition…

人机交互 · 计算机科学 2023-08-01 Omri Suissa , Avshalom Elmalech , Maayan Zhitomirsky-Geffet

Optical Character Recognition (OCR) of eighteenth-century printed texts remains challenging due to degraded print quality, archaic glyphs, and non-standardized orthography. Although transformer-based OCR systems and Vision-Language Models…

计算机视觉与模式识别 · 计算机科学 2026-02-17 Ari Vesalainen , Eetu Mäkelä , Laura Ruotsalainen , Mikko Tolonen

Digital humanities scholars increasingly use Large Language Models for historical document digitization, yet lack appropriate evaluation frameworks for LLM-based OCR. Traditional metrics fail to capture temporal biases and period-specific…

计算机视觉与模式识别 · 计算机科学 2025-10-09 Maria Levchenko

This paper addresses a major challenge to historical research on the 19th century. Large quantities of sources have become digitally available for the first time, while extraction techniques are lagging behind. Therefore, we researched…

计算机视觉与模式识别 · 计算机科学 2024-01-17 David Fleischhacker , Wolfgang Goederle , Roman Kern

For the bachelor project 2021 of Professor Lippert's research group, handwritten entries of historical patient records needed to be digitized using Optical Character Recognition (OCR) methods. Since the data will be used in the future, a…

计算机视觉与模式识别 · 计算机科学 2025-09-23 Martin Preiß

Optical Character Recognition (OCR) is a critical but error-prone stage in digital humanities text pipelines. While OCR correction improves usability for downstream NLP tasks, common workflows often overwrite intermediate decisions,…

人机交互 · 计算机科学 2026-05-07 Haoze Guo , Ziqi Wei

Linked Data is used in various fields as a new way of structuring and connecting data. Cultural heritage institutions have been using linked data to improve archival descriptions and facilitate the discovery of information. Most archival…

计算机视觉与模式识别 · 计算机科学 2023-11-28 Mariana Dias , Carla Teixeira Lopes

Extracting fine-grained OCR text from aged documents in diacritic languages remains challenging due to unexpected artifacts, time-induced degradation, and lack of datasets. While standalone spell correction approaches have been proposed,…

计算与语言 · 计算机科学 2025-02-28 Thao Do , Dinh Phu Tran , An Vo , Daeyoung Kim

Thousands of users consult digital archives daily, but the information they can access is unrepresentative of the diversity of documentary history. The sequence-to-sequence architecture typically used for optical character recognition (OCR)…

计算机视觉与模式识别 · 计算机科学 2024-07-29 Jacob Carlson , Tom Bryan , Melissa Dell

The digitization of historical folkloristic materials presents unique challenges due to diverse text layouts, varying print and handwriting styles, and linguistic variations. This study explores different optical character recognition (OCR)…

数字图书馆 · 计算机科学 2025-07-28 Octavian M. Machidon , Alina L. Machidon

Conventional Optical Character Recognition (OCR) systems are challenged by variant invoice layouts, handwritten text, and low-quality scans, which are often caused by strong template dependencies that restrict their flexibility across…

计算机视觉与模式识别 · 计算机科学 2026-01-09 Khushi Khanchandani , Advait Thakur , Akshita Shetty , Chaitravi Reddy , Ritisa Behera

This paper discusses how to successfully digitize large-scale historical micro-data by augmenting optical character recognition (OCR) engines with pre- and post-processing methods. Although OCR software has improved dramatically in recent…

计算机视觉与模式识别 · 计算机科学 2023-09-21 Sergio Correia , Stephan Luck

Word error rate of an ocr is often higher than its character error rate. This is especially true when ocrs are designed by recognizing characters. High word accuracies are critical to tasks like the creation of content in digital libraries…

计算机视觉与模式识别 · 计算机科学 2019-05-29 Deepayan Das , Jerin Philip , Minesh Mathew , C. V. Jawahar

This paper explores the application of synthetic data in the post-OCR domain on multiple fronts by conducting experiments to assess the impact of data volume, augmentation, and synthetic data generation methods on model performance.…

计算与语言 · 计算机科学 2024-08-14 Shuhao Guan , Derek Greene

Despite their cultural and historical significance, Black digital archives continue to be a structurally underrepresented area in AI research and infrastructure. This is especially evident in efforts to digitize historical Black newspapers,…

数字图书馆 · 计算机科学 2025-09-17 Fitsum Sileshi Beyene , Christopher L. Dancy
‹ 上一页 1 2 3 10 下一页 ›