English
Related papers

Related papers: Lipi Gnani - A Versatile OCR for Documents in any …

200 papers

The objective of the paper is to recognize handwritten samples of basic Bangla characters using Tesseract open source Optical Character Recognition (OCR) engine under Apache License 2.0. Handwritten data samples containing isolated Bangla…

Computer Vision and Pattern Recognition · Computer Science 2010-03-31 Sandip Rakshit , Debkumar Ghosal , Tanmoy Das , Subhrajit Dutta , Subhadip Basu

Optical character recognition (OCR) has advanced rapidly with the rise of vision-language models, yet evaluation has remained concentrated on a small cluster of high- and mid-resource scripts. We introduce GlotOCR Bench, a comprehensive…

Computation and Language · Computer Science 2026-04-15 Amir Hossein Kargaran , Nafiseh Nikeghbal , Jana Diesner , François Yvon , Hinrich Schütze

Word-level handwritten optical character recognition (OCR) remains a challenge for morphologically rich languages like Bangla. The complexity arises from the existence of a large number of alphabets, the presence of several diacritic forms,…

Computer Vision and Pattern Recognition · Computer Science 2022-05-24 Md. Ismail Hossain , Mohammed Rakib , Sabbir Mollah , Fuad Rahman , Nabeel Mohammed

Computer-Assisted Pronunciation Training (CAPT) has been extensively studied for English. However, there remains a critical gap in its application to Indian languages with a base of 1.5 billion speakers. Pronunciation tools tailored to…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-04 Arnav Rustagi , Satvik Bajpai , Nimrat Kaur , Siddharth Siddharth

This paper presents a novel approach to recognize Grantha, an ancient script in South India and converting it to Malayalam, a prevalent language in South India using online character recognition mechanism. The motivation behind this work…

Computer Vision and Pattern Recognition · Computer Science 2012-08-22 Sreeraj. M , Sumam Mary Idicula

We propose a post-OCR text correction approach for digitising texts in Romanised Sanskrit. Owing to the lack of resources our approach uses OCR models trained for other languages written in Roman. Currently, there exists no dataset…

Computation and Language · Computer Science 2018-09-10 Amrith Krishna , Bodhisattwa Prasad Majumder , Rajesh Shreedhar Bhat , Pawan Goyal

Kannudi is a reference editor for Kannada based on OPOK! and OHOK! principles, and domain knowledge. It introduces a method of input for Kannada, called OHOK!, that is, Ottu Haku Ottu Kodu! (apply pressure and give ottu). This is especially…

Human-Computer Interaction · Computer Science 2023-01-04 Vishweshwar V. Dixit

Retrieval-Augmented Generation (RAG) has become a popular technique for enhancing the reliability and utility of Large Language Models (LLMs) by grounding responses in external documents. Traditional RAG systems rely on Optical Character…

Computer Vision and Pattern Recognition · Computer Science 2025-05-12 Alexander Most , Joseph Winjum , Ayan Biswas , Shawn Jones , Nishath Rajiv Ranasinghe , Dan O'Malley , Manish Bhattarai

In this paper, we address the task of Optical Character Recognition(OCR) for the Telugu script. We present an end-to-end framework that segments the text image, classifies the characters and extracts lines using a language model. The…

Machine Learning · Statistics 2017-02-16 Rakesh Achanta , Trevor Hastie

Handwritten word recognition and spotting of low-resource scripts are difficult as sufficient training data is not available and it is often expensive for collecting data of such scripts. This paper presents a novel cross language platform…

Computer Vision and Pattern Recognition · Computer Science 2018-02-06 Ayan Kumar Bhunia , Partha Pratim Roy , Akash Mohta , Umapada Pal

Optical Character Recognition (OCR) is the process of extracting digitized text from images of scanned documents. While OCR systems have already matured in many languages, they still have shortcomings in cursive languages with overlapping…

Computer Vision and Pattern Recognition · Computer Science 2020-09-22 Hussein Osman , Karim Zaghw , Mostafa Hazem , Seifeldin Elsehely

This work presents an accuracy study of the open source OCR engine, Kraken, on the leading Arabic scholarly journal, al-Abhath. In contrast with other commercially available OCR engines, Kraken is shown to be capable of producing highly…

Computation and Language · Computer Science 2024-02-20 Benjamin Kiessling , Gennady Kurin , Matthew Thomas Miller , Kader Smail

The Optical Character Recognition (OCR) systems have been widely used in various of application scenarios, such as office automation (OA) systems, factory automations, online educations, map productions etc. However, OCR is still a…

Computer Vision and Pattern Recognition · Computer Science 2020-10-16 Yuning Du , Chenxia Li , Ruoyu Guo , Xiaoting Yin , Weiwei Liu , Jun Zhou , Yifan Bai , Zilin Yu , Yehua Yang , Qingqing Dang , Haoshuang Wang

The success rates of Optical Character Recognition (OCR) systems for printed Malayalam documents is quite impressive with the state of the art accuracy levels in the range of 85-95% for various. However for real applications, further…

Computation and Language · Computer Science 2012-05-09 Sajilal Divakaran

Traditional Optical Character Recognition (OCR) systems that generate text of highly inflectional Indic languages like Hindi tend to suffer from poor accuracy due to a wide alphabet set, compound characters and difficulty in segmenting…

Computation and Language · Computer Science 2020-12-15 Aditya Pal , Abhijit Mustafi

Automated language processing is central to the drive to enable facilitated referencing of increasingly available Sanskrit E texts. The first step towards processing Sanskrit text involves the handling of Sanskrit compound words that are an…

Computation and Language · Computer Science 2009-11-05 N. Rama , Meenakshi Lakshmanan

The objective of the paper is to recognize handwritten samples of lower case Roman script using Tesseract open source Optical Character Recognition (OCR) engine under Apache License 2.0. Handwritten data samples containing isolated and…

Computer Vision and Pattern Recognition · Computer Science 2010-03-31 Sandip Rakshit , Subhadip Basu

The Digital Corpus of Sanskrit records around 650,000 sentences along with their morphological and lexical tagging. But inconsistencies in morphological analysis, and in providing crucial information like the segmented word, urges the need…

Computation and Language · Computer Science 2020-05-15 Sriram Krishnan , Amba Kulkarni , Gérard Huet

The OpenITI team has achieved Optical Character Recognition (OCR) accuracy rates for classical Arabic-script texts in the high nineties. These numbers are based on our tests of seven different Arabic-script texts of varying quality and…

Computer Vision and Pattern Recognition · Computer Science 2017-03-29 Maxim Romanov , Matthew Thomas Miller , Sarah Bowen Savant , Benjamin Kiessling

This project undertakes the training and analysis of optical character recognition OCR methods applied to 10th century ancient Tamil inscriptions discovered on the walls of the Brihadeeswarar Temple.The chosen OCR methods include…

Computer Vision and Pattern Recognition · Computer Science 2024-05-29 Velmathi G , Shangavelan M , Harish D , Krithikshun M S