English
Related papers

Related papers: Chronicling Germany: An Annotated Historical Newsp…

200 papers

The use of propaganda has spiked on mainstream and social media, aiming to manipulate or mislead users. While efforts to automatically detect propaganda techniques in textual, visual, or multimodal content have increased, most of them…

Computation and Language · Computer Science 2024-02-28 Maram Hasanain , Fatema Ahmed , Firoj Alam

This paper describes a dataset containing small images of text from everyday scenes. The purpose of the dataset is to support the development of new automated systems that can detect and analyze text. Although much research has been devoted…

Computer Vision and Pattern Recognition · Computer Science 2016-10-21 Ahmed Ibrahim , A. Lynn Abbott , Mohamed E. Hussein

A common approach for improving OCR quality is a post-processing step based on models correcting misdetected characters and tokens. These models are typically trained on aligned pairs of OCR read text and their manually corrected…

Computation and Language · Computer Science 2019-06-27 Kai Hakala , Aleksi Vesanto , Niko Miekka , Tapio Salakoski , Filip Ginter

We introduce the Brno Mobile OCR Dataset (B-MOD) for document Optical Character Recognition from low-quality images captured by handheld mobile devices. While OCR of high-quality scanned documents is a mature field where many commercial…

Computer Vision and Pattern Recognition · Computer Science 2019-07-03 Martin Kišš , Michal Hradiš , Oldřich Kodym

Online propaganda poses a severe threat to the integrity of societies. However, existing datasets for detecting online propaganda have a key limitation: they were annotated using weak labels that can be noisy and even incorrect. To address…

Computation and Language · Computer Science 2024-11-27 Abdurahman Maarouf , Dominik Bär , Dominique Geissler , Stefan Feuerriegel

This paper describes a system prepared at Brno University of Technology for ICDAR 2021 Competition on Historical Document Classification, experiments leading to its design, and the main findings. The solved tasks include script and font…

Computer Vision and Pattern Recognition · Computer Science 2022-03-31 Martin Kišš , Jan Kohút , Karel Beneš , Michal Hradiš

Over the past few decades, large archives of paper-based documents such as books and newspapers have been digitized using Optical Character Recognition. This technology is error-prone, especially for historical documents. To correct OCR…

Computation and Language · Computer Science 2023-08-01 Omri Suissa , Avshalom Elmalech , Maayan Zhitomirsky-Geffet

Given the ubiquity of handwritten documents in human transactions, Optical Character Recognition (OCR) of documents have invaluable practical worth. Optical character recognition is a science that enables to translate various types of…

Computer Vision and Pattern Recognition · Computer Science 2020-01-03 Jamshed Memon , Maira Sami , Rizwan Ahmed Khan

In this paper, we develop machine learning techniques to identify unknown printers in early modern (c.~1500--1800) English printed books. Specifically, we focus on matching uniquely damaged character type-imprints in anonymously printed…

Computer Vision and Pattern Recognition · Computer Science 2023-06-16 Nikolai Vogler , Kartik Goyal , Kishore PV Reddy , Elizaveta Pertseva , Samuel V. Lemley , Christopher N. Warren , Max G'Sell , Taylor Berg-Kirkpatrick

Optical character recognition (OCR) is a fundamental problem in computer vision. Research studies have shown significant progress in classifying printed characters using deep learning-based methods and topologies. Among current algorithms,…

Computer Vision and Pattern Recognition · Computer Science 2018-04-12 Saman Sarraf

A foundational task for the digital analysis of documents is text line segmentation. However, automating this process with deep learning models is challenging because it requires large, annotated datasets that are often unavailable for…

Computer Vision and Pattern Recognition · Computer Science 2025-11-12 Rafael Sterzinger , Tingyu Lin , Robert Sablatnig

Page layout analysis is a fundamental step in document processing which enables to segment a page into regions of interest. With highly complex layouts and mixed scripts, scholarly commentaries are text-heavy documents which remain…

Information Retrieval · Computer Science 2022-12-29 Najem-Meyer Sven , Romanello Matteo

Latin has historically led the state-of-the-art in handwritten optical character recognition (OCR) research. Adapting existing systems from Latin to alpha-syllabary languages is particularly challenging due to a sharp contrast between their…

Computer Vision and Pattern Recognition · Computer Science 2022-06-30 Samiul Alam , Tahsin Reasat , Asif Shahriyar Sushmit , Sadi Mohammad Siddiquee , Fuad Rahman , Mahady Hasan , Ahmed Imtiaz Humayun

Automated headline generation for online news articles is not a trivial task - machine generated titles need to be grammatically correct, informative, capture attention and generate search traffic without being "click baits" or "fake news".…

Machine Learning · Computer Science 2021-07-26 Cristian Anastasiu , Hanna Behnke , Sarah Lück , Viktor Malesevic , Aamna Najmi , Javier Poveda-Panter

Nowadays, metadata information is often given by the authors themselves upon submission. However, a significant part of already existing research papers have missing or incomplete metadata information. German scientific papers come in a…

Information Retrieval · Computer Science 2021-11-11 Azeddine Bouabdallah , Jorge Gavilan , Jennifer Gerbl , Prayuth Patumcharoenpol

Optical character recognition (OCR) is a vital process that involves the extraction of handwritten or printed text from scanned or printed images, converting it into a format that can be understood and processed by machines. This enables…

Computer Vision and Pattern Recognition · Computer Science 2023-12-20 Mahmoud SalahEldin Kasem , Mohamed Mahmoud , Hyun-Soo Kang

We introduce the task of historical text summarisation, where documents in historical forms of a language are summarised in the corresponding modern language. This is a fundamentally important routine to historians and digital humanities…

Computation and Language · Computer Science 2022-01-25 Xutan Peng , Yi Zheng , Chenghua Lin , Advaith Siddharthan

This paper examines the current state-of-the-art of German text simplification, focusing on parallel and monolingual German corpora. It reviews neural language models for simplifying German texts and assesses their suitability for legal…

Computation and Language · Computer Science 2023-12-18 Thorben Schomacker , Michael Gille , Jörg von der Hülls , Marina Tropmann-Frick

The digitization of historical folkloristic materials presents unique challenges due to diverse text layouts, varying print and handwriting styles, and linguistic variations. This study explores different optical character recognition (OCR)…

Digital Libraries · Computer Science 2025-07-28 Octavian M. Machidon , Alina L. Machidon

The extraction of anglicisms (lexical borrowings from English) is relevant both for lexicographic purposes and for NLP downstream tasks. We introduce a corpus of European Spanish newspaper headlines annotated with anglicisms and a baseline…

Computation and Language · Computer Science 2020-04-08 Elena Álvarez-Mellado
‹ Prev 1 4 5 6 7 8 10 Next ›