English
Related papers

Related papers: Chronicling Germany: An Annotated Historical Newsp…

200 papers

Text line segmentation is one of the pre-stages of modern optical character recognition systems. The algorithmic approach proposed by this paper has been designed for this exact purpose. Its main characteristic is the combination of two…

Computer Vision and Pattern Recognition · Computer Science 2023-06-22 Pit Schneider

Introduction. We study effect of different quality optical character recognition in interactive information retrieval with a collection of one digitized historical Finnish newspaper. Method. This study is based on the simulated interactive…

Computation and Language · Computer Science 2022-06-02 Kimmo Kettunen , Heikki Keskustalo , Sanna Kumpulainen , Tuula Pääkkönen , Juha Rautiainen

A diversity of tasks use language models trained on semantic similarity data. While there are a variety of datasets that capture semantic similarity, they are either constructed from modern web data or are relatively small datasets created…

Computation and Language · Computer Science 2023-08-25 Emily Silcock , Melissa Dell

Thousands of users consult digital archives daily, but the information they can access is unrepresentative of the diversity of documentary history. The sequence-to-sequence architecture typically used for optical character recognition (OCR)…

Computer Vision and Pattern Recognition · Computer Science 2024-07-29 Jacob Carlson , Tom Bryan , Melissa Dell

Natural-language processing of historical documents is complicated by the abundance of variant spellings and lack of annotated data. A common approach is to normalize the spelling of historical words to modern forms. We explore the…

Computation and Language · Computer Science 2016-10-26 Marcel Bollmann , Anders Søgaard

This paper introduces a very challenging dataset of historic German documents and evaluates Fully Convolutional Neural Network (FCNN) based methods to locate handwritten annotations of any kind in these documents. The handwritten…

Computer Vision and Pattern Recognition · Computer Science 2018-12-07 Andreas Kölsch , Ashutosh Mishra , Saurabh Varshneya , Muhammad Zeshan Afzal , Marcus Liwicki

The National Library of Finland has digitized the historical newspapers published in Finland between 1771 and 1910. This collection contains approximately 1.95 million pages in Finnish and Swedish. Finnish part of the collection consists of…

Computation and Language · Computer Science 2019-10-18 Kimmo Kettunen

Data-driven analyses of biases in historical texts can help illuminate the origin and development of biases prevailing in modern society. However, digitised historical documents pose a challenge for NLP practitioners as these corpora suffer…

Computation and Language · Computer Science 2023-05-23 Nadav Borenstein , Karolina Stańczak , Thea Rolskov , Natália da Silva Perez , Natacha Klein Käfer , Isabelle Augenstein

Extraction of text regions and individual text lines from historic documents is necessary for automatic transcription. We propose extending a CNN-based text baseline detection system by adding line height and text block boundary predictions…

Computer Vision and Pattern Recognition · Computer Science 2021-02-24 Oldřich Kodym , Michal Hradiš

This Data Descriptor introduces the dataset Enevaeldens Nyheder Online (News during Absolutism Online). The Enevaeldens Nyheder Online (ENO) dataset provides a reconstruction of the contents of major newspapers in Denmark and Norway during…

Digital Libraries · Computer Science 2025-09-03 Johan Heinsen , Camilla Bøgeskov

Historical Document Processing is the process of digitizing written material from the past for future use by historians and other scholars. It incorporates algorithms and software tools from various subfields of computer science, including…

Computer Vision and Pattern Recognition · Computer Science 2020-09-14 James P. Philips , Nasseh Tabrizi

This paper addresses the challenge of automatically extracting attributes from news article web pages across multiple languages. Recent neural network models have shown high efficacy in extracting information from semi-structured web pages.…

Computation and Language · Computer Science 2025-02-05 Pavel Bedrin , Maksim Varlamov , Alexander Yatskov

We address the problem of segmenting and retrieving word images in collections of historical manuscripts given a text query. This is commonly referred to as "word spotting". To this end, we first propose an end-to-end trainable model based…

Computer Vision and Pattern Recognition · Computer Science 2020-04-02 Tomas Wilkinson , Jonas Lindström , Anders Brun

The growing volume of digitized historical texts requires effective semantic search using text embeddings. However, pre-trained multilingual models face challenges with historical content due to OCR noise and outdated spellings. This study…

Computation and Language · Computer Science 2025-03-14 Andrianos Michail , Corina Julia Raclé , Juri Opitz , Simon Clematide

The digitisation of historical documents has provided historians with unprecedented research opportunities. Yet, the conventional approach to analysing historical documents involves converting them from images to text using OCR, a process…

Computation and Language · Computer Science 2023-11-07 Nadav Borenstein , Phillip Rust , Desmond Elliott , Isabelle Augenstein

This research digitizes and analyzes the Leidse hoogleraren en lectoren 1575-1815 books written between 1983 and 1985, which contain biographic data about professors and curators of Leiden University. It addresses the central question: how…

Computation and Language · Computer Science 2026-01-01 Zahra Abedi , Richard M. K. van Dijk , Gijs Wijnholds , Tessa Verhoef

This paper investigates advertising practices in print newspapers across India using a novel data-driven approach. We develop a pipeline employing image processing and OCR techniques to extract articles and advertisements from digital…

Computers and Society · Computer Science 2025-05-19 N Harsha Vardhan , Ponnurangam Kumaraguru , Kiran Garimella

Most datasets in the field of document analysis utilize highly standardized labels, which, while simplifying specific tasks, often produce outputs that are not directly applicable to humanities research. In contrast, the Nuremberg…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Martin Mayr , Julian Krenz , Katharina Neumeier , Anna Bub , Simon Bürcky , Nina Brolich , Klaus Herbers , Mechthild Habermann , Peter Fleischmann , Andreas Maier , Vincent Christlein

This paper presents a pipeline with minimal human influence for scraping and detecting bias on college newspaper archives. This paper introduces a framework for scraping complex archive sites that automated tools fail to grab data from, and…

Computation and Language · Computer Science 2023-09-14 Adam M. Lehavi , William McCormack , Noah Kornfeld , Solomon Glazer

Effects of Optical Character Recognition (OCR) quality on historical information retrieval have so far been studied in data-oriented scenarios regarding the effectiveness of retrieval results. Such studies have either focused on the effects…

Information Retrieval · Computer Science 2022-08-12 Kimmo Kettunen , Heikki Keskustalo , Sanna Kumpulainen , Tuula Pääkkönen , Juha Rautiainen