English
Related papers

Related papers: PHD: Pixel-Based Language Modeling of Historical D…

200 papers

Data-driven analyses of biases in historical texts can help illuminate the origin and development of biases prevailing in modern society. However, digitised historical documents pose a challenge for NLP practitioners as these corpora suffer…

Computation and Language · Computer Science 2023-05-23 Nadav Borenstein , Karolina Stańczak , Thea Rolskov , Natália da Silva Perez , Natacha Klein Käfer , Isabelle Augenstein

Various algorithms have been proposed for dictionary learning. Among those for image processing, many use image patches to form dictionaries. This paper focuses on whole-image recovery from corrupted linear measurements. We address the open…

Computer Vision and Pattern Recognition · Computer Science 2014-08-19 Yangyang Xu , Wotao Yin

Optical character recognition (OCR) for historical documents is a complex procedure subject to a unique set of material issues, including inconsistencies in typefaces and low quality scanning. Consequently, even the most sophisticated OCR…

Computation and Language · Computer Science 2020-04-27 Alberto Poncelas , Mohammad Aboomar , Jan Buts , James Hadley , Andy Way

Automatic document content processing is affected by artifacts caused by the shape of the paper, non-uniform and diverse color of lighting conditions. Fully-supervised methods on real data are impossible due to the large amount of data…

Computer Vision and Pattern Recognition · Computer Science 2020-12-01 Sagnik Das , Hassan Ahmed Sial , Ke Ma , Ramon Baldrich , Maria Vanrell , Dimitris Samaras

This paper explores the application of synthetic data in the post-OCR domain on multiple fronts by conducting experiments to assess the impact of data volume, augmentation, and synthetic data generation methods on model performance.…

Computation and Language · Computer Science 2024-08-14 Shuhao Guan , Derek Greene

Historical maps are valuable resources that capture detailed geographical information from the past. However, these maps are typically available in printed formats, which are not conducive to modern computer-based analyses. Digitizing these…

Computer Vision and Pattern Recognition · Computer Science 2025-01-06 Yunshuang Yuan , Frank Thiemann , Monika Sester

As more historical texts are digitized, there is interest in applying natural language processing tools to these archives. However, the performance of these tools is often unsatisfactory, due to language change and genre differences.…

Computation and Language · Computer Science 2016-04-05 Yi Yang , Jacob Eisenstein

OCR errors are common in digitised historical archives significantly affecting their usability and value. Generative Language Models (LMs) have shown potential for correcting these errors using the context provided by the corrupted text and…

Computation and Language · Computer Science 2024-10-01 Jonathan Bourne

Detecting tampered text in document images is a challenging task due to data scarcity. To address this, previous work has attempted to generate tampered documents using rule-based methods. However, the resulting documents often suffer from…

Computer Vision and Pattern Recognition · Computer Science 2026-02-20 Mohamed Dhouib , Davide Buscaldi , Sonia Vanier , Aymen Shabou

Parliamentary proceedings represent a rich yet challenging resource for computational analysis, particularly when preserved only as scanned historical documents. Existing efforts to transcribe Italian parliamentary speeches have relied on…

Digital Libraries · Computer Science 2026-05-21 Luigi Curini , Alfio Ferrara , Giovanni Pagano , Sergio Picascia

Recently, the progress of learning-by-synthesis has proposed a training model for synthetic images, which can effectively reduce the cost of human and material resources. However, due to the different distribution of synthetic images…

Computer Vision and Pattern Recognition · Computer Science 2019-03-21 Tongtong Zhao , Yuxiao Yan , Jinjia Peng , Huibing Wang , Xianping Fu

The reconstruction of shredded documents consists in arranging the pieces of paper (shreds) in order to reassemble the original aspect of such documents. This task is particularly relevant for supporting forensic investigation as documents…

Computer Vision and Pattern Recognition · Computer Science 2020-04-30 Thiago M. Paixão , Rodrigo F. Berriel , Maria C. S. Boeres , Alessando L. Koerich , Claudine Badue , Alberto F. De Souza , Thiago Oliveira-Santos

The unprecedented success of image reconstruction approaches based on deep neural networks has revolutionised both the processing and the analysis paradigms in several applied disciplines. In the field of digital humanities, the task of…

Computer Vision and Pattern Recognition · Computer Science 2023-12-14 Fabio Merizzi , Perrine Saillard , Oceane Acquier , Elena Morotti , Elena Loli Piccolomini , Luca Calatroni , Rosa Maria Dessì

Image enhancement is an important image processing technique that processes images suitably for a specific application e.g. image editing. The conventional solutions of image enhancement are grouped into two categories which are spatial…

Computer Vision and Pattern Recognition · Computer Science 2016-09-14 Hui Li , Xiaomeng Wang , Weifeng Liu , Yanjiang Wang

Handwritten document-image binarization is a semantic segmentation process to differentiate ink pixels from background pixels. It is one of the essential steps towards character recognition, writer identification, and script-style evolution…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Maruf A. Dhali , Jan Willem de Wit , Lambert Schomaker

Historic variations of spelling poses a challenge for full-text search or natural language processing on historical digitized texts. To minimize the gap between the historic orthography and contemporary spelling, usually an automatic…

Computation and Language · Computer Science 2025-02-26 Anton Ehrmanntraut

Efficient categorization of historical documents is crucial for fields such as genealogy, legal research, and historical scholarship, where manual classification is impractical for large collections due to its labor-intensive and…

Computer Vision and Pattern Recognition · Computer Science 2024-05-27 Taylor Archibald , Tony Martinez

The indexing and searching of historical documents have garnered attention in recent years due to massive digitization efforts of important collections worldwide. Pure textual search in these corpora is a problem since optical character…

Information Retrieval · Computer Science 2020-04-23 Taivanbat Badamdorj , Adiel Ben-Shalom , Nachum Dershowitz , Lior Wolf

Pseudocode in a scholarly paper provides a concise way to express the algorithms implemented therein. Pseudocode can also be thought of as an intermediary representation that helps bridge the gap between programming languages and natural…

Information Retrieval · Computer Science 2024-06-10 Levent Toksoz , Gang Tan , C. Lee Giles

Handwritten Text Recognition (HTR) in free-layout pages is a challenging image understanding task that can provide a relevant boost to the digitization of handwritten documents and reuse of their content. The task becomes even more…

Computer Vision and Pattern Recognition · Computer Science 2022-08-18 Silvia Cascianelli , Marcella Cornia , Lorenzo Baraldi , Rita Cucchiara