English
Related papers

Related papers: Jochre 3 and the Yiddish OCR corpus

200 papers

In this paper, we address the task of Optical Character Recognition(OCR) for the Telugu script. We present an end-to-end framework that segments the text image, classifies the characters and extracts lines using a language model. The…

Machine Learning · Statistics 2017-02-16 Rakesh Achanta , Trevor Hastie

Diacritization of Arabic text is both an interesting and a challenging problem at the same time with various applications ranging from speech synthesis to helping students learning the Arabic language. Like many other tasks or problems in…

Computation and Language · Computer Science 2019-05-07 Ali Fadel , Ibraheem Tuffaha , Bara' Al-Jawarneh , Mahmoud Al-Ayyoub

We present PashtoCorp, a 1.25-billion-word corpus for Pashto, a language spoken by 60 million people that remains severely underrepresented in NLP. The corpus is assembled from 39 sources spanning seven HuggingFace datasets and 32…

Computation and Language · Computer Science 2026-03-18 Hanif Rahman

Optical character recognition (OCR) technology has been widely used in various scenes, as shown in Figure 1. Designing a practical OCR system is still a meaningful but challenging task. In previous work, considering the efficiency and…

Computer Vision and Pattern Recognition · Computer Science 2022-06-15 Chenxia Li , Weiwei Liu , Ruoyu Guo , Xiaoting Yin , Kaitao Jiang , Yongkun Du , Yuning Du , Lingfeng Zhu , Baohua Lai , Xiaoguang Hu , Dianhai Yu , Yanjun Ma

The objective of the paper is to recognize handwritten samples of lower case Roman script using Tesseract open source Optical Character Recognition (OCR) engine under Apache License 2.0. Handwritten data samples containing isolated and…

Computer Vision and Pattern Recognition · Computer Science 2010-03-31 Sandip Rakshit , Subhadip Basu

We present OCR-Quality, a comprehensive human-annotated dataset designed for evaluating and developing OCR quality assessment methods. The dataset consists of 1,000 PDF pages converted to PNG images at 300 DPI, sampled from diverse…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Yulong Zhang

In order to apply Optical Character Recognition (OCR) to historical printings of Latin script fully automatically, we report on our efforts to construct a widely-applicable polyfont recognition model yielding text with a Character Error…

Computer Vision and Pattern Recognition · Computer Science 2021-06-16 Christian Reul , Christoph Wick , Maximilian Nöth , Andreas Büttner , Maximilian Wehner , Uwe Springmann

Billions of public domain documents remain trapped in hard copy or lack an accurate digitization. Modern natural language processing methods cannot be used to index, retrieve, and summarize their texts; conduct computational textual…

Computer Vision and Pattern Recognition · Computer Science 2023-10-17 Tom Bryan , Jacob Carlson , Abhishek Arora , Melissa Dell

Kurdish is a less-resourced language consisting of different dialects written in various scripts. Approximately 30 million people in different countries speak the language. The lack of corpora is one of the main obstacles in Kurdish…

Computation and Language · Computer Science 2019-09-26 Roshna Omer Abdulrahman , Hossein Hassani , Sina Ahmadi

Open-weight LLMs have been released by frontier labs; however, sovereign Large Language Models (for languages other than English) remain low in supply yet high in demand. Training large language models (LLMs) for low-resource languages such…

Computation and Language · Computer Science 2026-02-03 Shaltiel Shmidman , Avi Shmidman , Amir DN Cohen , Moshe Koppel

This research is the second phase in a series of investigations on developing an Optical Character Recognition (OCR) of Arabic historical documents and examining how different modeling procedures interact with the problem. The first…

Computer Vision and Pattern Recognition · Computer Science 2022-08-30 Aly Mostafa , Omar Mohamed , Ali Ashraf , Ahmed Elbehery , Salma Jamal , Anas Salah , Amr S. Ghoneim

Over the past few decades, large archives of paper-based documents such as books and newspapers have been digitized using Optical Character Recognition. This technology is error-prone, especially for historical documents. To correct OCR…

Computation and Language · Computer Science 2023-08-01 Omri Suissa , Avshalom Elmalech , Maayan Zhitomirsky-Geffet

Intensive research has been done on optical character recognition ocr and a large number of articles have been published on this topic during the last few decades. Many commercial OCR systems are now available in the market, but most of…

Computer Vision and Pattern Recognition · Computer Science 2016-09-08 K. Indira , S. Sethu Selvi

Over the past few decades, large archives of paper-based historical documents, such as books and newspapers, have been digitized using the Optical Character Recognition (OCR) technology. Unfortunately, this broadly used technology is…

Computation and Language · Computer Science 2023-08-01 Omri Suissa , Maayan Zhitomirsky-Geffet , Avshalom Elmalech

Optical character recognition (OCR) has advanced rapidly with the rise of vision-language models, yet evaluation has remained concentrated on a small cluster of high- and mid-resource scripts. We introduce GlotOCR Bench, a comprehensive…

Computation and Language · Computer Science 2026-04-15 Amir Hossein Kargaran , Nafiseh Nikeghbal , Jana Diesner , François Yvon , Hinrich Schütze

Substantial amounts of work are required to clean large collections of digitized books for NLP analysis, both because of the presence of errors in the scanned text and the presence of duplicate volumes in the corpora. In this paper, we…

Computation and Language · Computer Science 2021-10-25 Allen Kim , Charuta Pethe , Naoya Inoue , Steve Skiena

We present the Knesset Corpus, a corpus of Hebrew parliamentary proceedings containing over 30 million sentences (over 384 million tokens) from all the (plenary and committee) protocols held in the Israeli parliament between 1998 and 2022.…

Computation and Language · Computer Science 2025-06-02 Gili Goldin , Nick Howell , Noam Ordan , Ella Rabinovich , Shuly Wintner

This contribution to a special issue on "Computer-aided processing of intertextuality" in ancient texts will illustrate how using digital tools to interact with the Hebrew Bible offers new promising perspectives for visualizing the texts…

Computation and Language · Computer Science 2017-10-25 Nicolai Winther-Nielsen

Offensive language detection has been well studied in many languages, but it is lagging behind in low-resource languages, such as Hebrew. In this paper, we present a new offensive language corpus in Hebrew. A total of 15,881 tweets were…

Computation and Language · Computer Science 2023-09-07 Nagham Hamad , Mustafa Jarrar , Mohammad Khalilia , Nadim Nashif

In this paper, we introduce the Chinese corpus from CLUE organization, CLUECorpus2020, a large-scale corpus that can be used directly for self-supervised learning such as pre-training of a language model, or language generation. It has 100G…

Computation and Language · Computer Science 2020-03-06 Liang Xu , Xuanwei Zhang , Qianqian Dong