English
Related papers

Related papers: OCR Synthetic Benchmark Dataset for Indic Language…

200 papers

Multimodal research has predominantly focused on single-image reasoning, with limited exploration of multi-image scenarios. Recent models have sought to enhance multi-image understanding through large-scale pretraining on interleaved…

Computation and Language · Computer Science 2026-03-26 Shaharukh Khan , Ali Faraz , Abhinav Ravi , Mohd Nauman , Mohd Sarfraz , Akshat Patidar , Raja Kolla , Chandra Khatri , Shubham Agarwal

Cross-Lingual SynthDocs is a large-scale synthetic corpus designed to address the scarcity of Arabic resources for Optical Character Recognition (OCR) and Document Understanding (DU). The dataset comprises over 2.5 million of samples,…

Computation and Language · Computer Science 2025-11-10 Haneen Al-Homoud , Asma Ibrahim , Murtadha Al-Jubran , Fahad Al-Otaibi , Yazeed Al-Harbi , Daulet Toibazar , Kesen Wang , Pedro J. Moreno

To assist in the development of machine learning methods for automated classification of spectroscopic data, we have generated a universal synthetic dataset that can be used for model validation. This dataset contains artificial spectra…

Machine Learning · Computer Science 2022-06-15 Jan Schuetzke , Nathan J. Szymanski , Markus Reischl

New deep-learning architectures are created every year, achieving state-of-the-art results in image recognition and leading to the belief that, in a few years, complex tasks such as sign language translation will be considerably easier,…

Computer Vision and Pattern Recognition · Computer Science 2021-04-05 Alvaro Leandro Cavalcante Carneiro , Lucas de Brito Silva , Denis Henrique Pinheiro Salvadeo

Industrial Retrieval-Augmented Generation (RAG) systems depend on optical character recognition (OCR) to transform visual documents into text. Existing OCR benchmarks rely on character-level metrics, which inadequately measure downstream…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Lin Sun , Wang Dexian , Jingang Huang , Linglin Zhang , Change Jia , Zhengwei Cheng , Xiangzheng Zhang

Handwritten character recognition is a challenging research in the field of document image analysis over many decades due to numerous reasons such as large writing styles variation, inherent noise in data, expansive applications it offers,…

Computer Vision and Pattern Recognition · Computer Science 2021-07-21 Noushath Shaffi , Faizal Hajamohideen

Natural Language Generation (NLG) for non-English languages is hampered by the scarcity of datasets in these languages. In this paper, we present the IndicNLG Benchmark, a collection of datasets for benchmarking NLG for 11 Indic languages.…

Computation and Language · Computer Science 2022-10-28 Aman Kumar , Himani Shrotriya , Prachi Sahu , Raj Dabre , Ratish Puduppully , Anoop Kunchukuttan , Amogh Mishra , Mitesh M. Khapra , Pratyush Kumar

In this paper, we introduce Neural Information Retrieval resources for 11 widely spoken Indian Languages (Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Oriya, Punjabi, Tamil, and Telugu) from two major Indian language…

Information Retrieval · Computer Science 2023-12-18 Saiful Haq , Ashutosh Sharma , Pushpak Bhattacharyya

Training a conventional automatic speech recognition (ASR) system to support multiple languages is challenging because the sub-word unit, lexicon and word inventories are typically language specific. In contrast, sequence-to-sequence models…

Audio and Speech Processing · Electrical Eng. & Systems 2018-02-16 Shubham Toshniwal , Tara N. Sainath , Ron J. Weiss , Bo Li , Pedro Moreno , Eugene Weinstein , Kanishka Rao

We aim to investigate the performance of current OCR systems on low resource languages and low resource scripts. We introduce and make publicly available a novel benchmark, OCR4MT, consisting of real and synthetic data, enriched with noise,…

Computation and Language · Computer Science 2022-03-15 Oana Ignat , Jean Maillard , Vishrav Chaudhary , Francisco Guzmán

This paper presents a printed Bengali and English text OCR system developed by us using a single hidden BLSTM-CTC architecture having 128 units. Here, we did not use any peephole connection and dropout in the BLSTM, which helped us in…

Computer Vision and Pattern Recognition · Computer Science 2019-08-26 Debabrata Paul , Bidyut Baran Chaudhuri

Optical character recognition (OCR) is a process of converting analogue documents into digital using document images. Currently, many commercial and non-commercial OCR systems exist for both handwritten and printed copies for different…

Computer Vision and Pattern Recognition · Computer Science 2021-05-11 Farisa Benta Safir , Abu Quwsar Ohi , M. F. Mridha , Muhammad Mostafa Monowar , Md. Abdul Hamid

Reading scene text, that is, text appearing in images, has numerous application areas, including assistive technology, search, and e-commerce. Although scene text recognition in English has advanced significantly and is often considered…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Anik De , Abhirama Subramanyam Penamakuri , Rajeev Yadav , Aditya Rathore , Harshiv Shah , Devesh Sharma , Sagar Agarwal , Pravin Kumar , Anand Mishra

Automatic License Plate Recognition is a frequent research topic due to its wide-ranging practical applications. While recent studies use synthetic images to improve License Plate Recognition (LPR) results, there remain several limitations…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Rayson Laroca , Valter Estevam , Gladston J. P. Moreira , Rodrigo Minetto , David Menotti

OCR errors are common in digitised historical archives significantly affecting their usability and value. Generative Language Models (LMs) have shown potential for correcting these errors using the context provided by the corrupted text and…

Computation and Language · Computer Science 2024-10-01 Jonathan Bourne

Thousands of users consult digital archives daily, but the information they can access is unrepresentative of the diversity of documentary history. The sequence-to-sequence architecture typically used for optical character recognition (OCR)…

Computer Vision and Pattern Recognition · Computer Science 2024-07-29 Jacob Carlson , Tom Bryan , Melissa Dell

Urdu, spoken by over 250 million people, remains critically under-served in multimodal and vision-language research. The absence of large-scale, high-quality datasets has limited the development of Urdu-capable systems and reinforced biases…

Computer Vision and Pattern Recognition · Computer Science 2025-09-12 Umair Hassan

Scoring the Optical Character Recognition (OCR) capabilities of Large Multimodal Models (LMMs) has witnessed growing interest. Existing benchmarks have highlighted the impressive performance of LMMs in text recognition; however, their…

Indian Sign Language has limited resources for developing machine learning and data-driven approaches for automated language processing. Though text/audio-based language processing techniques have shown colossal research interest and…

Computation and Language · Computer Science 2024-07-09 Abhinav Joshi , Romit Mohanty , Mounika Kanakanti , Andesha Mangla , Sudeep Choudhary , Monali Barbate , Ashutosh Modi

Arabic Optical Character Recognition (OCR) is essential for converting vast amounts of Arabic print media into digital formats. However, training modern OCR models, especially powerful vision-language models, is hampered by the lack of…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Omer Nacar , Yasser Al-Habashi , Serry Sibaee , Adel Ammar , Wadii Boulila