English
Related papers

Related papers: American Stories: A Large-Scale Structured Text Da…

200 papers

Chronicling America is a product of the National Digital Newspaper Program, a partnership between the Library of Congress and the National Endowment for the Humanities to digitize historic newspapers. Over 16 million pages of historic…

In the U.S. historically, local newspapers drew their content largely from newswires like the Associated Press. Historians argue that newswires played a pivotal role in creating a national identity and shared understanding of the world, but…

Computation and Language · Computer Science 2024-06-17 Emily Silcock , Abhishek Arora , Luca D'Amico-Wong , Melissa Dell

Digitization of newspapers is of interest for many reasons including preservation of history, accessibility and search ability, etc. While digitization of documents such as scientific articles and magazines is prevalent in literature, one…

Computer Vision and Pattern Recognition · Computer Science 2022-02-04 Wenzhen Zhu , Negin Sokhandan , Guang Yang , Sujitha Martin , Suchitra Sathyanarayana

A diversity of tasks use language models trained on semantic similarity data. While there are a variety of datasets that capture semantic similarity, they are either constructed from modern web data or are relatively small datasets created…

Computation and Language · Computer Science 2023-08-25 Emily Silcock , Melissa Dell

The correct detection of dense article layout and the recognition of characters in historical newspaper pages remains a challenging requirement for Natural Language Processing (NLP) and machine learning applications on historical newspapers…

Digital Libraries · Computer Science 2025-06-17 Christian Schultze , Niklas Kerkfeld , Kara Kuebart , Princilia Weber , Moritz Wolter , Felix Selgert

Most tools for accessing digitized historical newspapers emphasize relatively simple search; but, as increasing numbers of digitized historical newspapers and other historical resources become available we can consider much richer modes of…

Digital Libraries · Computer Science 2015-02-16 Robert B. Allen

Newspapers are documents made of news item and informative articles. They are not meant to be red iteratively: the reader can pick his items in any order he fancies. Ignoring this structural property, most digitized newspaper archives only…

Information Retrieval · Computer Science 2012-10-04 Thomas Palfray , David Hébert , Stéphane Nicolas , Pierrick Tranouez , Thierry Paquet

This paper describes the construction of a large-scale corpus of historical wire articles from U.S. Southern newspapers, spanning 1960-1975 and covering multiple wire services (e.g., Associated Press, United Press International, Newspaper…

Computation and Language · Computer Science 2026-01-19 Michael McRae

Unsupervised discovery of stories with correlated news articles in real-time helps people digest massive news streams without expensive human annotations. A common approach of the existing studies for unsupervised online story discovery is…

Information Retrieval · Computer Science 2023-05-05 Susik Yoon , Dongha Lee , Yunyi Zhang , Jiawei Han

Web articles such as Wikipedia serve as one of the major sources of knowledge dissemination and online learning. However, their in-depth information--often in a dense text format--may not be suitable for mobile browsing, even in a…

Human-Computer Interaction · Computer Science 2023-10-05 Daniel Nkemelu , Peggy Chi , Daniel Castro Chin , Krishna Srinivasan , Irfan Essa

The New York Public Library is participating in the Chronicling America initiative to develop an online searchable database of historically significant newspaper articles. Microfilm copies of the newspapers are scanned and high resolution…

Text segmentation, the task of dividing a document into sections, is often a prerequisite for performing additional natural language processing tasks. Existing text segmentation methods have typically been developed and tested using clean,…

Computer Vision and Pattern Recognition · Computer Science 2023-12-21 Carol Anderson , Phil Crone

Content analysis of news stories (whether manual or automatic) is a cornerstone of the communication studies field. However, much research is conducted at the level of individual news articles, despite the fact that news events (especially…

Social and Information Networks · Computer Science 2024-10-31 Tom Nicholls , Jonathan Bright

The massive amounts of digitized historical documents acquired over the last decades naturally lend themselves to automatic processing and exploration. Research work seeking to automatically process facsimiles and extract information…

Computer Vision and Pattern Recognition · Computer Science 2023-06-22 Raphaël Barman , Maud Ehrmann , Simon Clematide , Sofia Ares Oliveira , Frédéric Kaplan

Despite their cultural and historical significance, Black digital archives continue to be a structurally underrepresented area in AI research and infrastructure. This is especially evident in efforts to digitize historical Black newspapers,…

Digital Libraries · Computer Science 2025-09-17 Fitsum Sileshi Beyene , Christopher L. Dancy

There is an overwhelming number of news articles published every day around the globe. Following the evolution of a news-story is a difficult task given that there is no such mechanism available to track back in time to study the diffusion…

Information Retrieval · Computer Science 2017-12-22 Roberto Camacho Barranco , Arnold P. Boedihardjo , M. Shahriar Hossain

Following a particular news story online is an important but difficult task, as the relevant information is often scattered across different domains/sources (e.g., news articles, blogs, comments, tweets), presented in various formats and…

Computation and Language · Computer Science 2018-08-20 Bichen Shi , Thanh-Binh Le , Neil Hurley , Georgiana Ifrim

Traditional information retrieval is primarily concerned with finding relevant information from large datasets without imposing a structure within the retrieved pieces of data. However, structuring information in the form of…

Information Retrieval · Computer Science 2025-03-21 Fausto German , Brian Keith , Chris North

There is a huge amount of historical documents in libraries and in various National Archives that have not been exploited electronically. Although automatic reading of complete pages remains, in most cases, a long-term objective, tasks such…

Computer Vision and Pattern Recognition · Computer Science 2007-05-23 Laurence Likforman-Sulem , Abderrazak Zahour , Bruno Taconet

With rapidly evolving media narratives, it has become increasingly critical to not just extract narratives from a given corpus but rather investigate, how they develop over time. While popular narrative extraction methods such as Large…

Computation and Language · Computer Science 2025-06-26 Kai-Robin Lange , Tobias Schmidt , Matthias Reccius , Henrik Müller , Michael Roos , Carsten Jentsch
‹ Prev 1 2 3 10 Next ›