English
Related papers

Related papers: Logical segmentation for article extraction in dig…

200 papers

Despite their cultural and historical significance, Black digital archives continue to be a structurally underrepresented area in AI research and infrastructure. This is especially evident in efforts to digitize historical Black newspapers,…

Digital Libraries · Computer Science 2025-09-17 Fitsum Sileshi Beyene , Christopher L. Dancy

Pagination - the process of determining where to break an article across pages in a multi-article layout is a common layout challenge for most commercially printed newspapers and magazines. To date, no one has created an algorithm that…

Computation and Language · Computer Science 2014-04-15 Joshua Hailpern , Niranjan Damera Venkata , Marina Danilevsky

Neutrality is difficult to achieve and, in politics, subjective. Traditional media typically adopt an editorial line that can be used by their potential readers as an indicator of the media bias. Several platforms currently rate news…

Computation and Language · Computer Science 2023-10-26 Cristina España-Bonet

This paper introduces a new way for text-line extraction by integrating deep-learning based pre-classification and state-of-the-art segmentation methods. Text-line extraction in complex handwritten documents poses a significant challenge,…

Computer Vision and Pattern Recognition · Computer Science 2019-07-02 Michele Alberti , Lars Vögtlin , Vinaychandran Pondenkandath , Mathias Seuret , Rolf Ingold , Marcus Liwicki

Text line segmentation is one of the key steps in historical document understanding. It is challenging due to the variety of fonts, contents, writing styles and the quality of documents that have degraded through the years. In this paper,…

Computer Vision and Pattern Recognition · Computer Science 2022-10-24 Mélodie Boillet , Christopher Kermorvant , Thierry Paquet

The segmentation of complex images into semantic regions has seen a growing interest these last years with the advent of Deep Learning. Until recently, most existing methods for Historical Document Analysis focused on the visual appearance…

Computer Vision and Pattern Recognition · Computer Science 2021-11-03 Mélodie Boillet , Martin Maarand , Thierry Paquet , Christopher Kermorvant

In the age of information overload, content management for online news articles relies on efficient summarization to enhance accessibility and user engagement. This article addresses the challenge of extractive text summarization by…

Machine Learning · Computer Science 2025-09-22 Sajib Biswas , Milon Biswas , Arunima Mandal , Fatema Tabassum Liza , Joy Sarker

In this paper, we bring a new way of digesting news content by introducing the task of segmenting a news article into multiple sections and generating the corresponding summary to each section. We make two contributions towards this new…

Computation and Language · Computer Science 2021-10-18 Yang Liu , Chenguang Zhu , Michael Zeng

This paper introduces Fundus, a user-friendly news scraper that enables users to obtain millions of high-quality news articles with just a few lines of code. Unlike existing news scrapers, we use manually crafted, bespoke content extractors…

Computation and Language · Computer Science 2024-06-25 Max Dallabetta , Conrad Dobberstein , Adrian Breiding , Alan Akbik

There is an overwhelming number of news articles published every day around the globe. Following the evolution of a news-story is a difficult task given that there is no such mechanism available to track back in time to study the diffusion…

Information Retrieval · Computer Science 2017-12-22 Roberto Camacho Barranco , Arnold P. Boedihardjo , M. Shahriar Hossain

This paper presents a computational method of analysis that draws from machine learning, library science, and literary studies to map the visual layouts of multi-ethnic newspapers from the late 19th and early 20th century United States.…

Computer Vision and Pattern Recognition · Computer Science 2021-09-07 Benjamin Charles Germain Lee , Joshua Ortiz Baco , Sarah H. Salter , Jim Casey

Information extraction (IE) from unstructured documents remains a critical challenge in data processing pipelines. Traditional optical character recognition (OCR) methods and conventional parsing engines demonstrate limited effectiveness…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Aditya Parikh

Within the past few decades we have witnessed digital revolution, which moved scholarly communication to electronic media and also resulted in a substantial increase in its volume. Nowadays keeping track with the latest scientific…

Digital Libraries · Computer Science 2017-10-30 Dominika Tkaczyk

Document Layout Analysis is a fundamental step in Handwritten Text Processing systems, from the extraction of the text lines to the type of zone it belongs to. We present a system based on artificial neural networks which is able to…

Computer Vision and Pattern Recognition · Computer Science 2018-12-13 Lorenzo Quirós

The explosion in the amount of news and journalistic content being generated across the globe, coupled with extended and instantaneous access to information through online media, makes it difficult and time-consuming to monitor news…

Computation and Language · Computer Science 2018-08-06 M. Tarik Altuncu , Sophia N. Yaliraki , Mauricio Barahona

Automatic extraction of procedural graphs from documents creates a low-cost way for users to easily understand a complex procedure by skimming visual graphs. Despite the progress in recent studies, it remains unanswered: whether the…

Computation and Language · Computer Science 2024-08-09 Weihong Du , Wenrui Liao , Hongru Liang , Wenqiang Lei

Digital archiving is becoming widespread owing to its effectiveness in protecting valuable books and providing knowledge to many people electronically. In this paper, we propose a novel approach to leverage digital archives for machine…

Computer Vision and Pattern Recognition · Computer Science 2023-10-04 Yamato Okamoto , Haruto Toyonaga , Yoshihisa Ijiri , Hirokatsu Kataoka

Tables present summarized and structured information to the reader, which makes table structure extraction an important part of document understanding applications. However, table structure identification is a hard problem not only because…

Computer Vision and Pattern Recognition · Computer Science 2020-02-07 Saqib Ali Khan , Syed Muhammad Daniyal Khalid , Muhammad Ali Shahzad , Faisal Shafait

Automatically detecting discourse segments is an important preliminary step towards full discourse parsing. Previous research on discourse segmentation have relied on the assumption that elementary discourse units (EDUs) in a document…

Computation and Language · Computer Science 2010-03-30 Stergos Afantenos , Pascal Denis , Philippe Muller , Laurence Danlos

Digitized archives contain and preserve the knowledge of generations of scholars in millions of documents. The size of these archives calls for automatic analysis since a manual analysis by specialists is often too expensive. In this paper,…

Computer Vision and Pattern Recognition · Computer Science 2020-11-05 Christian Bartz , Hendrik Rätz , Christoph Meinel