English
Related papers

Related papers: Cleaning Dirty Books: Post-OCR Processing for Prev…

200 papers

The great amount of information that can be stored in electronic media is growing up daily. Many of them is got mainly by typing, such as the huge of information obtained from web 2.0 sites; or scaned and processing by an Optical Character…

Computation and Language · Computer Science 2021-12-06 Wulfrano A. Luna-Ramírez , Carlos R. Jaimez-González

We report upon the results of a research and prototype building project \emph{Worldly~OCR} dedicated to developing new, more accurate image-to-text conversion software for several languages and writing systems. These include the cursive…

Computer Vision and Pattern Recognition · Computer Science 2020-05-19 Marek Rychlik , Dwight Nwaigwe , Yan Han , Dylan Murphy

Natural Language Processing (NLP) research is increasingly focusing on the use of Large Language Models (LLMs), with some of the most popular ones being either fully or partially closed-source. The lack of access to model details,…

Computation and Language · Computer Science 2024-02-23 Simone Balloccu , Patrícia Schmidtová , Mateusz Lango , Ondřej Dušek

Large models have recently played a dominant role in natural language processing and multimodal vision-language learning. However, their effectiveness in text-related visual tasks remains relatively unexplored. In this paper, we conducted a…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Yuliang Liu , Zhang Li , Mingxin Huang , Biao Yang , Wenwen Yu , Chunyuan Li , Xucheng Yin , Cheng-lin Liu , Lianwen Jin , Xiang Bai

Over the past decade, machine learning methods have given us driverless cars, voice recognition, effective web search, and a much better understanding of the human genome. Machine learning is so common today that it is used dozens of times…

Computer Vision and Pattern Recognition · Computer Science 2021-06-22 Omer Aydin

Detecting and recognizing text in natural scene images is a challenging, yet not completely solved task. In re- cent years several new systems that try to solve at least one of the two sub-tasks (text detection and text recognition) have…

Computer Vision and Pattern Recognition · Computer Science 2017-07-28 Christian Bartz , Haojin Yang , Christoph Meinel

This study utilizes machine learning algorithms to analyze and organize knowledge in the field of algorithmic trading. By filtering a dataset of 136 million research papers, we identified 14,342 relevant articles published between 1956 and…

Statistical Finance · Quantitative Finance 2024-11-11 Stanisław Łaniewski , Robert Ślepaczuk

Accurate transcription of handwritten mathematics is crucial for educational AI systems, yet current benchmarks fail to evaluate this capability properly. Most prior studies focus on single-line expressions and rely on lexical metrics such…

Computers and Society · Computer Science 2026-05-27 Jin Seong , Wencke Liermann , Minho Kim , Jong-hun Shin , Soojong Lim

In this report, I present a deep learning approach to conduct a natural language processing (hereafter NLP) binary classification task for analyzing financial-fraud texts. First, I searched for regulatory announcements and enforcement…

Computation and Language · Computer Science 2023-08-09 Qiuru Li

Detecting security vulnerabilities in software before they are exploited has been a challenging problem for decades. Traditional code analysis methods have been proposed, but are often ineffective and inefficient. In this work, we model…

Cryptography and Security · Computer Science 2021-05-07 Noah Ziems , Shaoen Wu

Removing noise from scanned pages is a vital step before their submission to the optical character recognition (OCR) system. Most available image denoising methods are supervised where the pairs of noisy/clean pages are required. However,…

Computer Vision and Pattern Recognition · Computer Science 2021-10-11 Mehrdad J Gangeh , Marcin Plata , Hamid Motahari , Nigel P Duffy

Purpose: To develop a deep learning approach to digitally-stain optical coherence tomography (OCT) images of the optic nerve head (ONH). Methods: A horizontal B-scan was acquired through the center of the ONH using OCT (Spectralis) for 1…

This paper introduces a novel approach to post-Optical Character Recognition Correction (POC) for handwritten Cyrillic text, addressing a significant gap in current research methodologies. This gap is due to the lack of large text corporas…

Computer Vision and Pattern Recognition · Computer Science 2023-11-28 Evgenii Davydkin , Aleksandr Markelov , Egor Iuldashev , Anton Dudkin , Ivan Krivorotov

State-of-the-art models can perform well in controlled environments, but they often struggle when presented with out-of-distribution (OOD) examples, making OOD detection a critical component of NLP systems. In this paper, we focus on…

Computation and Language · Computer Science 2023-07-17 Mateusz Baran , Joanna Baran , Mateusz Wójcik , Maciej Zięba , Adam Gonczarek

In existing splicing forgery datasets, the insufficient semantic variety of spliced regions causes trained detection models to overfit semantic features rather than learn genuine splicing traces. Meanwhile, the lack of a reasonable…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Jiaming Liang , Yuwan Xue , Haowei Liu , Zhenqi Dai , Yu Liao , Rui Wang , Weihao Jiang , Yaping Liu , Zhikun Chen , Guoxiao Liu , Bo Liu , Xiuli Bi

Historical documents frequently suffer from damage and inconsistencies, including missing or illegible text resulting from issues such as holes, ink problems, and storage damage. These missing portions or gaps are referred to as lacunae. In…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Jaydeep Borkar , David A. Smith

The digitization of multi-domain retail billing documents remains a challenging task due to variability in scan quality, layout heterogeneity, and domain diversity across commercial sectors. This paper proposes and benchmarks an…

Computer Vision and Pattern Recognition · Computer Science 2026-04-29 Vijaysinh Gaikwad

With the growing interest in large language models, the need for evaluating the quality of machine text compared to reference (typically human-generated) text has become focal attention. Most recent works focus either on task-specific…

We describe a system used by the NASA Astrophysics Data System to identify bibliographic references obtained from scanned article pages by OCR methods with records in a bibliographic database. We analyze the process generating the noisy…

Digital Libraries · Computer Science 2007-05-23 Markus Demleitner , Michael Kurtz , Alberto Accomazzi , Günther Eichhorn , Carolyn S. Grant , Steven S. Murray