English
Related papers

Related papers: Chronicling Germany: An Annotated Historical Newsp…

200 papers

Existing full text datasets of U.S. public domain newspapers do not recognize the often complex layouts of newspaper scans, and as a result the digitized content scrambles texts from articles, headlines, captions, advertisements, and other…

Computation and Language · Computer Science 2023-08-25 Melissa Dell , Jacob Carlson , Tom Bryan , Emily Silcock , Abhishek Arora , Zejiang Shen , Luca D'Amico-Wong , Quan Le , Pablo Querubin , Leander Heldring

Chronicling America is a product of the National Digital Newspaper Program, a partnership between the Library of Congress and the National Endowment for the Humanities to digitize historic newspapers. Over 16 million pages of historic…

Digitization of newspapers is of interest for many reasons including preservation of history, accessibility and search ability, etc. While digitization of documents such as scientific articles and magazines is prevalent in literature, one…

Computer Vision and Pattern Recognition · Computer Science 2022-02-04 Wenzhen Zhu , Negin Sokhandan , Guang Yang , Sujitha Martin , Suchitra Sathyanarayana

Historic variations of spelling poses a challenge for full-text search or natural language processing on historical digitized texts. To minimize the gap between the historic orthography and contemporary spelling, usually an automatic…

Computation and Language · Computer Science 2025-02-26 Anton Ehrmanntraut

Open-access corpora are essential for advancing natural language processing (NLP) and computational social science (CSS). However, large-scale resources for German remain limited, restricting research on linguistic trends and societal…

Computation and Language · Computer Science 2025-06-09 Stefanie Urchs , Veronika Thurner , Matthias Aßenmacher , Christian Heumann , Stephanie Thiemichen

The New York Public Library is participating in the Chronicling America initiative to develop an online searchable database of historically significant newspaper articles. Microfilm copies of the newspapers are scanned and high resolution…

Digital libraries oftentimes provide access to historical newspaper archives via keyword-based search. Historical figures and their roles are particularly interesting cognitive access points in historical research. Structuring and…

Digital Libraries · Computer Science 2023-07-19 Hermann Kroll , Christin Katharina Kreutz , Mirjam Cuper , Bill Matthias Thang , Wolf-Tilo Balke

When digitizing historical archives, it is necessary to search for the faces of celebrities and ordinary people, especially in newspapers, link them to the surrounding text, and make them searchable. Existing face detectors on datasets of…

Computer Vision and Pattern Recognition · Computer Science 2025-07-16 Marek Vaško , Adam Herout , Michal Hradiš

Optical character recognition (OCR) is crucial for a deeper access to historical collections. OCR needs to account for orthographic variations, typefaces, or language evolution (i.e., new letters, word spellings), as the main source of…

Computation and Language · Computer Science 2021-02-02 Lijun Lyu , Maria Koutraki , Martin Krickl , Besnik Fetahu

Newspapers are documents made of news item and informative articles. They are not meant to be red iteratively: the reader can pick his items in any order he fancies. Ignoring this structural property, most digitized newspaper archives only…

Information Retrieval · Computer Science 2012-10-04 Thomas Palfray , David Hébert , Stéphane Nicolas , Pierrick Tranouez , Thierry Paquet

Most tools for accessing digitized historical newspapers emphasize relatively simple search; but, as increasing numbers of digitized historical newspapers and other historical resources become available we can consider much richer modes of…

Digital Libraries · Computer Science 2015-02-16 Robert B. Allen

The indexing and searching of historical documents have garnered attention in recent years due to massive digitization efforts of important collections worldwide. Pure textual search in these corpora is a problem since optical character…

Information Retrieval · Computer Science 2020-04-23 Taivanbat Badamdorj , Adiel Ben-Shalom , Nachum Dershowitz , Lior Wolf

One important and particularly challenging step in the optical character recognition (OCR) of historical documents with complex layouts, such as newspapers, is the separation of text from non-text content (e.g. page borders or…

Computer Vision and Pattern Recognition · Computer Science 2020-04-17 Bernhard Liebl , Manuel Burghardt

NLP methods can aid historians in analyzing textual materials in greater volumes than manually feasible. Developing such methods poses substantial challenges though. First, acquiring large, annotated historical datasets is difficult, as…

Computation and Language · Computer Science 2023-05-19 Nadav Borenstein , Natalia da Silva Perez , Isabelle Augenstein

Despite their cultural and historical significance, Black digital archives continue to be a structurally underrepresented area in AI research and infrastructure. This is especially evident in efforts to digitize historical Black newspapers,…

Digital Libraries · Computer Science 2025-09-17 Fitsum Sileshi Beyene , Christopher L. Dancy

We introduce the AnnoPage Dataset, a novel collection of 7,550 pages from historical documents, primarily in Czech and German, spanning from 1485 to the present, focusing on the late 19th and early 20th centuries. The dataset is designed to…

Computer Vision and Pattern Recognition · Computer Science 2025-07-17 Martin Kišš , Michal Hradiš , Martina Dvořáková , Václav Jiroušek , Filip Kersch

Extracting metadata from scientific papers can be considered a solved problem in NLP due to the high accuracy of state-of-the-art methods. However, this does not apply to German scientific publications, which have a variety of styles and…

Information Retrieval · Computer Science 2021-06-15 Zeyd Boukhers , Nada Beili , Timo Hartmann , Prantik Goswami , Muhammad Arslan Zafar

Historical newspapers are a source of research for the human and social sciences. However, these image collections are difficult to read by machine due to the low quality of the print, the lack of standardization of the pages in addition to…

Information Retrieval · Computer Science 2020-02-21 José E. B. Maia , Gildácio J. de A. Sá

The massive amounts of digitized historical documents acquired over the last decades naturally lend themselves to automatic processing and exploration. Research work seeking to automatically process facsimiles and extract information…

Computer Vision and Pattern Recognition · Computer Science 2023-06-22 Raphaël Barman , Maud Ehrmann , Simon Clematide , Sofia Ares Oliveira , Frédéric Kaplan

This paper provides the first thorough documentation of a high quality digitization process applied to an early printed book from the incunabulum period (1450-1500). The entire OCR related workflow including preprocessing, layout analysis…

Computer Vision and Pattern Recognition · Computer Science 2017-01-26 Christian Reul , Marco Dittrich , Martin Gruner
‹ Prev 1 2 3 10 Next ›