中文
相关论文

相关论文: Digitizing Historical Balance Sheet Data: A Practi…

200 篇论文

Optical character recognition (OCR) is crucial for a deeper access to historical collections. OCR needs to account for orthographic variations, typefaces, or language evolution (i.e., new letters, word spellings), as the main source of…

计算与语言 · 计算机科学 2021-02-02 Lijun Lyu , Maria Koutraki , Martin Krickl , Besnik Fetahu

Optical Character Recognition (OCR) continues to face accuracy challenges that impact subsequent applications. To address these errors, we explore the utility of OCR confidence scores for enhancing post-OCR error detection. Our study…

计算机视觉与模式识别 · 计算机科学 2024-09-09 Arthur Hemmer , Mickaël Coustaty , Nicola Bartolo , Jean-Marc Ogier

Bahnar, a minority language spoken across Vietnam, Cambodia, and Laos, faces significant preservation challenges due to limited research and data availability. This study addresses the critical need for accurate digitization of Bahnar…

计算与语言 · 计算机科学 2026-01-07 Phat Tran , Phuoc Pham , Hung Trinh , Tho Quan

Information Extraction from visually rich documents is a challenging task that has gained a lot of attention in recent years due to its importance in several document-control based applications and its widespread commercial value. The…

计算机视觉与模式识别 · 计算机科学 2023-05-03 Mohamed Dhouib , Ghassen Bettaieb , Aymen Shabou

Digitization of newspapers is of interest for many reasons including preservation of history, accessibility and search ability, etc. While digitization of documents such as scientific articles and magazines is prevalent in literature, one…

计算机视觉与模式识别 · 计算机科学 2022-02-04 Wenzhen Zhu , Negin Sokhandan , Guang Yang , Sujitha Martin , Suchitra Sathyanarayana

Conventional optical character recognition (OCR) techniques segmented each character and then recognized. This made them prone to error in character segmentation, and devoid of context to exploit language models. Advances in sequence to…

计算机视觉与模式识别 · 计算机科学 2025-09-01 Shashank Vempati , Nishit Anand , Gaurav Talebailkar , Arpan Garai , Chetan Arora

Optical Character Recognition (OCR) systems often introduce errors when transcribing historical documents, leaving room for post-correction to improve text quality. This study evaluates the use of open-weight LLMs for OCR error correction…

计算与语言 · 计算机科学 2025-02-04 Jenna Kanerva , Cassandra Ledins , Siiri Käpyaho , Filip Ginter

Billions of public domain documents remain trapped in hard copy or lack an accurate digitization. Modern natural language processing methods cannot be used to index, retrieve, and summarize their texts; conduct computational textual…

计算机视觉与模式识别 · 计算机科学 2023-10-17 Tom Bryan , Jacob Carlson , Abhishek Arora , Melissa Dell

This paper presents a complete Optical Character Recognition (OCR) system for camera captured image/graphics embedded textual documents for handheld devices. At first, text regions are extracted and skew corrected. Then, these regions are…

计算机视觉与模式识别 · 计算机科学 2011-09-16 Ayatullah Faruk Mollah , Nabamita Majumder , Subhadip Basu , Mita Nasipuri

Linked Data is used in various fields as a new way of structuring and connecting data. Cultural heritage institutions have been using linked data to improve archival descriptions and facilitate the discovery of information. Most archival…

计算机视觉与模式识别 · 计算机科学 2023-11-28 Mariana Dias , Carla Teixeira Lopes

Current OCR systems are based on deep learning models trained on large amounts of data. Although they have shown some ability to generalize to unseen data, especially in detection tasks, they can struggle with recognizing low-quality data.…

Contrary to popular belief, Optical Character Recognition (OCR) remains a challenging problem when text occurs in unconstrained environments, like natural scenes, due to geometrical distortions, complex backgrounds, and diverse fonts. In…

计算机视觉与模式识别 · 计算机科学 2019-06-06 Marcin Namysl , Iuliu Konya

At a time when the quantity of - more or less freely - available data is increasing significantly, thanks to digital corpora, editions or libraries, the development of data mining tools or deep learning methods allows researchers to build a…

计算机视觉与模式识别 · 计算机科学 2019-04-29 Jean-Baptiste Camps , Gilles Guilhem Couffignal

The digitisation of historical documents has provided historians with unprecedented research opportunities. Yet, the conventional approach to analysing historical documents involves converting them from images to text using OCR, a process…

计算与语言 · 计算机科学 2023-11-07 Nadav Borenstein , Phillip Rust , Desmond Elliott , Isabelle Augenstein

This research digitizes and analyzes the Leidse hoogleraren en lectoren 1575-1815 books written between 1983 and 1985, which contain biographic data about professors and curators of Leiden University. It addresses the central question: how…

计算与语言 · 计算机科学 2026-01-01 Zahra Abedi , Richard M. K. van Dijk , Gijs Wijnholds , Tessa Verhoef

Optical character recognition (OCR) methods have been applied to diverse tasks, e.g., street view text recognition and document analysis. Recently, zero-shot OCR has piqued the interest of the research community because it considers a…

计算机视觉与模式识别 · 计算机科学 2023-08-02 Xiaolei Diao , Daqian Shi , Jian Li , Lida Shi , Mingzhe Yue , Ruihua Qi , Chuntao Li , Hao Xu

There has been recent interest in improving optical character recognition (OCR) for endangered languages, particularly because a large number of documents and books in these languages are not in machine-readable formats. The performance of…

计算与语言 · 计算机科学 2023-02-28 Shruti Rijhwani , Daisy Rosenblum , Michayla King , Antonios Anastasopoulos , Graham Neubig

Text line segmentation is one of the pre-stages of modern optical character recognition systems. The algorithmic approach proposed by this paper has been designed for this exact purpose. Its main characteristic is the combination of two…

计算机视觉与模式识别 · 计算机科学 2023-06-22 Pit Schneider

The digitization of multi-domain retail billing documents remains a challenging task due to variability in scan quality, layout heterogeneity, and domain diversity across commercial sectors. This paper proposes and benchmarks an…

计算机视觉与模式识别 · 计算机科学 2026-04-29 Vijaysinh Gaikwad

We propose a post-OCR text correction approach for digitising texts in Romanised Sanskrit. Owing to the lack of resources our approach uses OCR models trained for other languages written in Roman. Currently, there exists no dataset…

计算与语言 · 计算机科学 2018-09-10 Amrith Krishna , Bodhisattwa Prasad Majumder , Rajesh Shreedhar Bhat , Pawan Goyal