English
Related papers

Related papers: Jochre 3 and the Yiddish OCR corpus

200 papers

Thousands of users consult digital archives daily, but the information they can access is unrepresentative of the diversity of documentary history. The sequence-to-sequence architecture typically used for optical character recognition (OCR)…

Computer Vision and Pattern Recognition · Computer Science 2024-07-29 Jacob Carlson , Tom Bryan , Melissa Dell

In recent years, the field of Handwritten Text Recognition (HTR) has seen the emergence of various new models, each claiming to perform competitively better than the other in specific scenarios. However, making a fair comparison of these…

Computer Vision and Pattern Recognition · Computer Science 2025-02-11 Badri Vishal Kasuba , Dhruv Kudale , Venkatapathy Subramanian , Parag Chaudhuri , Ganesh Ramakrishnan

Handwriting recognition is one of the active and challenging areas of research in the field of image processing and pattern recognition. It has many applications that include: a reading aid for visual impairment, automated reading and…

Computer Vision and Pattern Recognition · Computer Science 2022-10-26 Rebin M. Ahmed , Tarik A. Rashid , Polla Fattah , Abeer Alsadoon , Nebojsa Bacanin , Seyedali Mirjalili , S. Vimal , Amit Chhabra

At a time when the quantity of - more or less freely - available data is increasing significantly, thanks to digital corpora, editions or libraries, the development of data mining tools or deep learning methods allows researchers to build a…

Computer Vision and Pattern Recognition · Computer Science 2019-04-29 Jean-Baptiste Camps , Gilles Guilhem Couffignal

Information Extraction from visually rich documents is a challenging task that has gained a lot of attention in recent years due to its importance in several document-control based applications and its widespread commercial value. The…

Computer Vision and Pattern Recognition · Computer Science 2023-05-03 Mohamed Dhouib , Ghassen Bettaieb , Aymen Shabou

UNESCO has classified 2500 out of 7000 languages spoken worldwide as endangered. Attrition of a language leads to loss of traditional wisdom, folk literature, and the essence of the community that uses it. It is therefore imperative to…

Computation and Language · Computer Science 2025-10-14 Prawaal Sharma , Poonam Goyal , Vidisha Sharma , Navneet Goyal

Scientific articles published prior to the "age of digitization" in the late 1990s contain figures which are "trapped" within their scanned pages. While progress to extract figures and their captions has been made, there is currently no…

Instrumentation and Methods for Astrophysics · Physics 2022-09-13 J. P. Naiman , Peter K. G. Williams , Alyssa Goodman

Deep learning based natural language processing model is proven powerful, but need large-scale dataset. Due to the significant gap between the real-world tasks and existing Chinese corpus, in this paper, we introduce a large-scale corpus of…

Computation and Language · Computer Science 2018-11-27 Jianyu Zhao , Zhuoran Ji

In this paper, we present a new Russian and Kazakh database (with about 95% of Russian and 5% of Kazakh words/sentences respectively) for offline handwriting recognition. A few pre-processing and segmentation procedures have been developed…

Computer Vision and Pattern Recognition · Computer Science 2021-08-31 Daniyar Nurseitov , Kairat Bostanbekov , Daniyar Kurmankhojayev , Anel Alimova , Abdelrahman Abdallah

Machine translation has been a major motivation of development in natural language processing. Despite the burgeoning achievements in creating more efficient machine translation systems thanks to deep learning methods, parallel corpora have…

Computation and Language · Computer Science 2020-10-06 Sina Ahmadi , Hossein Hassani , Daban Q. Jaff

This paper presents a comprehensive evaluation of the Optical Character Recognition (OCR) capabilities of the recently released GPT-4V(ision), a Large Multimodal Model (LMM). We assess the model's performance across a range of OCR tasks,…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Yongxin Shi , Dezhi Peng , Wenhui Liao , Zening Lin , Xinhong Chen , Chongyu Liu , Yuyi Zhang , Lianwen Jin

Developing a Bangla OCR requires bunch of algorithm and methods. There were many effort went on for developing a Bangla OCR. But all of them failed to provide an error free Bangla OCR. Each of them has some lacking. We discussed about the…

Computer Vision and Pattern Recognition · Computer Science 2012-04-06 Farjana Yeasmin Omee , Shiam Shabbir Himel , Md. Abu Naser Bikas

We introduce a large and diverse Czech corpus annotated for grammatical error correction (GEC) with the aim to contribute to the still scarce data resources in this domain for languages other than English. The Grammar Error Correction…

Computation and Language · Computer Science 2022-04-22 Jakub Náplava , Milan Straka , Jana Straková , Alexandr Rosen

Logo detection has been gaining considerable attention because of its wide range of applications in the multimedia field, such as copyright infringement detection, brand visibility monitoring, and product brand management on social media.…

Computer Vision and Pattern Recognition · Computer Science 2020-08-13 Jing Wang , Weiqing Min , Sujuan Hou , Shengnan Ma , Yuanjie Zheng , Shuqiang Jiang

Somali is a Cushitic language of the Horn of Africa with ~25 million speakers, yet no documented dedicated Somali pretraining corpus with a companion tokenizer and language-identification benchmark has been publicly released. Existing…

Computation and Language · Computer Science 2026-05-19 Khalid Yusuf Dahir

Independent of established data centers, and partly for my own research, since 1989 I have been collecting the tabular data from over 2600 articles concerned with radio sources and extragalactic objects in general. Optical character…

Instrumentation and Methods for Astrophysics · Physics 2009-11-13 Heinz Andernach

Thanks to improvements in machine learning techniques, including deep learning, speech synthesis is becoming a machine learning task. To accelerate speech synthesis research, we are developing Japanese voice corpora reasonably accessible…

Retrieving accurate details from documents is a crucial task, especially when handling a combination of scanned images and native digital formats. This document presents a combined framework for text extraction that merges Optical Character…

Computer Vision and Pattern Recognition · Computer Science 2025-06-16 Rasha Sinha , Rekha B S

Judeo-Arabic refers to Arabic variants historically spoken by Jewish communities across the Arab world, primarily during the Middle Ages. Unlike standard Arabic, it is written in Hebrew script by Jewish writers and for Jewish audiences.…

Computation and Language · Computer Science 2026-01-30 Juan Moreno Gonzalez , Bashar Alhafni , Nizar Habash

We present a new release of the Czech-English parallel corpus CzEng 2.0 consisting of over 2 billion words (2 "gigawords") in each language. The corpus contains document-level information and is filtered with several techniques to lower the…

Computation and Language · Computer Science 2020-07-08 Tom Kocmi , Martin Popel , Ondrej Bojar
‹ Prev 1 8 9 10 Next ›