English
Related papers

Related papers: Optical character recognition quality affects perc…

200 papers

Effects of Optical Character Recognition (OCR) quality on historical information retrieval have so far been studied in data-oriented scenarios regarding the effectiveness of retrieval results. Such studies have either focused on the effects…

Information Retrieval · Computer Science 2022-08-12 Kimmo Kettunen , Heikki Keskustalo , Sanna Kumpulainen , Tuula Pääkkönen , Juha Rautiainen

The National Library of Finland has digitized the historical newspapers published in Finland between 1771 and 1910. This collection contains approximately 1.95 million pages in Finnish and Swedish. Finnish part of the collection consists of…

Computation and Language · Computer Science 2019-10-18 Kimmo Kettunen

We implemented a high-performance optical character recognition model for classical handwritten documents using data augmentation with highly variable cropping within the document region. Optical character recognition in handwritten…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Joonmo Ahna , Taehong Jang , Quan Fengnyu , Hyungil Lee , Jaehyuk Lee , Sojung Lucia Kim

Named Entity Recognition (NER), search, classification and tagging of names and name like frequent informational elements in texts, has become a standard information extraction procedure for textual data. NER has been applied to many types…

Computation and Language · Computer Science 2016-11-10 Kimmo Kettunen , Eetu Mäkelä , Teemu Ruokolainen , Juha Kuokkala , Laura Löfberg

Today, most newsreaders read the online version of news articles rather than traditional paper-based newspapers. Also, news media publishers rely heavily on the income generated from subscriptions and website visits made by newsreaders.…

Information Retrieval · Computer Science 2020-04-21 Amin Omidvar , Hossein Poormodheji , Aijun An , Gordon Edall

Most tools for accessing digitized historical newspapers emphasize relatively simple search; but, as increasing numbers of digitized historical newspapers and other historical resources become available we can consider much richer modes of…

Digital Libraries · Computer Science 2015-02-16 Robert B. Allen

The correct detection of dense article layout and the recognition of characters in historical newspaper pages remains a challenging requirement for Natural Language Processing (NLP) and machine learning applications on historical newspapers…

Digital Libraries · Computer Science 2025-06-17 Christian Schultze , Niklas Kerkfeld , Kara Kuebart , Princilia Weber , Moritz Wolter , Felix Selgert

Historical newspapers are a source of research for the human and social sciences. However, these image collections are difficult to read by machine due to the low quality of the print, the lack of standardization of the pages in addition to…

Information Retrieval · Computer Science 2020-02-21 José E. B. Maia , Gildácio J. de A. Sá

Digital libraries oftentimes provide access to historical newspaper archives via keyword-based search. Historical figures and their roles are particularly interesting cognitive access points in historical research. Structuring and…

Digital Libraries · Computer Science 2023-07-19 Hermann Kroll , Christin Katharina Kreutz , Mirjam Cuper , Bill Matthias Thang , Wolf-Tilo Balke

Digital information exchange enables quick creation and sharing of information and thus changes existing habits. Social media is becoming the main source of news for end-users replacing traditional media. This also enables the proliferation…

Information Theory · Computer Science 2021-10-04 Aljaž Zrnec , Marko Poženel , Dejan Lavbič

Thousands of users consult digital archives daily, but the information they can access is unrepresentative of the diversity of documentary history. The sequence-to-sequence architecture typically used for optical character recognition (OCR)…

Computer Vision and Pattern Recognition · Computer Science 2024-07-29 Jacob Carlson , Tom Bryan , Melissa Dell

The New York Public Library is participating in the Chronicling America initiative to develop an online searchable database of historically significant newspaper articles. Microfilm copies of the newspapers are scanned and high resolution…

A common approach for improving OCR quality is a post-processing step based on models correcting misdetected characters and tokens. These models are typically trained on aligned pairs of OCR read text and their manually corrected…

Computation and Language · Computer Science 2019-06-27 Kai Hakala , Aleksi Vesanto , Niko Miekka , Tapio Salakoski , Filip Ginter

Purpose: The purpose of this paper is to investigate the impact of cooperative principle on the information quality (IQ) by making objects more relevant for consumer needs, in particular case Wikipedia articles for students.…

Computers and Society · Computer Science 2018-07-11 Miloš Fidler , Dejan Lavbič

The automatic quality assessment of self-media online articles is an urgent and new issue, which is of great value to the online recommendation and search. Different from traditional and well-formed articles, self-media online articles are…

Computation and Language · Computer Science 2020-08-14 Yiru Wang , Shen Huang , Gongfu Li , Qiang Deng , Dongliang Liao , Pengda Si , Yujiu Yang , Jin Xu

The indexing and searching of historical documents have garnered attention in recent years due to massive digitization efforts of important collections worldwide. Pure textual search in these corpora is a problem since optical character…

Information Retrieval · Computer Science 2020-04-23 Taivanbat Badamdorj , Adiel Ben-Shalom , Nachum Dershowitz , Lior Wolf

Given the ubiquity of handwritten documents in human transactions, Optical Character Recognition (OCR) of documents have invaluable practical worth. Optical character recognition is a science that enables to translate various types of…

Computer Vision and Pattern Recognition · Computer Science 2020-01-03 Jamshed Memon , Maira Sami , Rizwan Ahmed Khan

Currently, the quality of a search engine is often determined using so-called topical relevance, i.e., the match between the user intent (expressed as a query) and the content of the document. In this work we want to draw attention to two…

Information Retrieval · Computer Science 2015-01-27 Aleksandr Chuklin , Maarten de Rijke

Recognition of historical documents is a challenging problem due to the noised, damaged characters and background. However, in Japanese historical documents, not only contains the mentioned problems, pre-modern Japanese characters were…

Computer Vision and Pattern Recognition · Computer Science 2019-05-15 Anh Duc Le , Tarin Clanuwat , Asanobu Kitamoto

Rapid increase of digitized document give birth to high demand of document image retrieval. While conventional document image retrieval approaches depend on complex OCR-based text recognition and text similarity detection, this paper…

Computer Vision and Pattern Recognition · Computer Science 2017-09-04 Mao Tan , Si-Ping Yuan , Yong-Xin Su
‹ Prev 1 2 3 10 Next ›