English
Related papers

Related papers: Aspect-Driven Structuring of Historical Dutch News…

200 papers

Most tools for accessing digitized historical newspapers emphasize relatively simple search; but, as increasing numbers of digitized historical newspapers and other historical resources become available we can consider much richer modes of…

Digital Libraries · Computer Science 2015-02-16 Robert B. Allen

The New York Public Library is participating in the Chronicling America initiative to develop an online searchable database of historically significant newspaper articles. Microfilm copies of the newspapers are scanned and high resolution…

Existing full text datasets of U.S. public domain newspapers do not recognize the often complex layouts of newspaper scans, and as a result the digitized content scrambles texts from articles, headlines, captions, advertisements, and other…

Computation and Language · Computer Science 2023-08-25 Melissa Dell , Jacob Carlson , Tom Bryan , Emily Silcock , Abhishek Arora , Zejiang Shen , Luca D'Amico-Wong , Quan Le , Pablo Querubin , Leander Heldring

Many libraries offer free access to digitised historical newspapers via user interfaces. After an initial period of search and filter options as the only features, the availability of more advanced tools and the desire for more options…

Digital Libraries · Computer Science 2020-06-05 Eva Pfanzelter , Sarah Oberbichler , Jani Marjanen , Pierre-Carl Langlais , Stefan Hechl

The correct detection of dense article layout and the recognition of characters in historical newspaper pages remains a challenging requirement for Natural Language Processing (NLP) and machine learning applications on historical newspapers…

Digital Libraries · Computer Science 2025-06-17 Christian Schultze , Niklas Kerkfeld , Kara Kuebart , Princilia Weber , Moritz Wolter , Felix Selgert

Newspapers are documents made of news item and informative articles. They are not meant to be red iteratively: the reader can pick his items in any order he fancies. Ignoring this structural property, most digitized newspaper archives only…

Information Retrieval · Computer Science 2012-10-04 Thomas Palfray , David Hébert , Stéphane Nicolas , Pierrick Tranouez , Thierry Paquet

Providing effective access paths to content is a key task in digital libraries. Oftentimes, such access paths are realized through advanced query languages, which, on the one hand, users may find challenging to learn or use, and on the…

Digital Libraries · Computer Science 2023-05-01 Hermann Kroll , Christin Katharina Kreutz , Pascal Sackhoff , Wolf-Tilo Balke

Designing keyword-based access paths is a common practice in digital libraries. They are easy to use and accepted by users and come with moderate costs for content providers. However, users usually have to break down the search into pieces…

Digital Libraries · Computer Science 2022-08-23 Hermann Kroll , Niklas Mainzer , Wolf-Tilo Balke

NLP methods can aid historians in analyzing textual materials in greater volumes than manually feasible. Developing such methods poses substantial challenges though. First, acquiring large, annotated historical datasets is difficult, as…

Computation and Language · Computer Science 2023-05-19 Nadav Borenstein , Natalia da Silva Perez , Isabelle Augenstein

The task of organizing and clustering multilingual news articles for media monitoring is essential to follow news stories in real time. Most approaches to this task focus on high-resource languages (mostly English), with low-resource…

Computation and Language · Computer Science 2022-04-29 João Santos , Afonso Mendes , Sebastião Miranda

Archived collections of documents (like newspaper archives) serve as important information sources for historians, journalists, sociologists and other interested parties. Semantic Layers over such digital archives allow describing and…

Information Retrieval · Computer Science 2022-10-19 Pavlos Fafalios , Vaibhav Kasturia , Wolfgang Nejdl

With the enrichment of literature resources, researchers are facing the growing problem of information explosion and knowledge overload. To help scholars retrieve literature and acquire knowledge successfully, clarifying the semantic…

Computation and Language · Computer Science 2021-12-03 Bowen Ma , Chengzhi Zhang , Yuzhuo Wang , Sanhong Deng

Longitudinal corpora like newspaper archives are of immense value to historical research, and time as an important factor for historians strongly influences their search behaviour in these archives. While searching for articles published…

Information Retrieval · Computer Science 2018-10-25 Jaspreet Singh , Wolfgang Nejdl , Avishek Anand

Modern news aggregators do the hard work of organizing a large news stream, creating collections for a given news story with tens of source options. This paper shows that navigating large source collections for a news story can be…

Human-Computer Interaction · Computer Science 2023-02-20 Philippe Laban , Chien-Sheng Wu , Lidiya Murakhovs'ka , Xiang 'Anthony' Chen , Caiming Xiong

Thousands of users consult digital archives daily, but the information they can access is unrepresentative of the diversity of documentary history. The sequence-to-sequence architecture typically used for optical character recognition (OCR)…

Computer Vision and Pattern Recognition · Computer Science 2024-07-29 Jacob Carlson , Tom Bryan , Melissa Dell

Historians and archivists often find and analyze the occurrences of query words in newspaper archives, to help answer fundamental questions about society. But much work in text analytics focuses on helping people investigate other textual…

Human-Computer Interaction · Computer Science 2022-04-12 Abram Handler , Narges Mahyar , Brendan O'Connor

Digitization of newspapers is of interest for many reasons including preservation of history, accessibility and search ability, etc. While digitization of documents such as scientific articles and magazines is prevalent in literature, one…

Computer Vision and Pattern Recognition · Computer Science 2022-02-04 Wenzhen Zhu , Negin Sokhandan , Guang Yang , Sujitha Martin , Suchitra Sathyanarayana

Purpose: Advanced usage of Web Analytics tools allows to capture the content of user queries. Despite their relevant nature, the manual analysis of large volumes of user queries is problematic. This paper demonstrates the potential of using…

Information Retrieval · Computer Science 2017-09-25 Anne Chardonnens , Ettore Rizza , Mathias Coeckelbergs , Seth van Hooland

Digital libraries maintain extensive collections of knowledge and need to provide effective access paths for their users. For instance, PubPharm, the specialized information service for Pharmacy in Germany, provides and develops access…

Information Retrieval · Computer Science 2025-09-15 Hermann Kroll , Pascal Sackhoff , Bill Matthias Thang , Christin Katharina Kreutz , Wolf-Tilo Balke

This paper proposes OCR++, an open-source framework designed for a variety of information extraction tasks from scholarly articles including metadata (title, author names, affiliation and e-mail), structure (section headings and body text,…

‹ Prev 1 2 3 10 Next ›