中文
相关论文

相关论文: Lipi Gnani - A Versatile OCR for Documents in any …

200 篇论文

Automatic Cognate Detection (ACD) is a challenging task which has been utilized to help NLP applications like Machine Translation, Information Retrieval and Computational Phylogenetics. Unidentified cognate pairs can pose a challenge to…

计算与语言 · 计算机科学 2022-01-03 Diptesh Kanojia , Kevin Patel , Pushpak Bhattacharyya , Malhar Kulkarni , Gholamreza Haffari

In this paper, I present our work on DeepRAG, a specialized embedding model we built specifically for Hindi language in RAG systems. While LLMs have gotten really good at generating text, their performance in retrieval tasks still depends…

计算与语言 · 计算机科学 2025-03-12 Nandakishor M

This research investigates biases in text-to-image (TTI) models for the Indic languages widely spoken across India. It evaluates and compares the generative performance and cultural relevance of leading TTI models in these languages against…

计算与语言 · 计算机科学 2024-08-02 Surbhi Mittal , Arnav Sudan , Mayank Vatsa , Richa Singh , Tamar Glaser , Tal Hassner

In this research work, we have proposed an algorithm based on supervised learning methodology to extract the root forms of the Bengali verbs using the grammatical rules proposed by Panini [1] in Ashtadhyayi. This methodology can be applied…

计算与语言 · 计算机科学 2020-04-02 Arijit Das , Tapas Halder , Diganta Saha

Character recognition techniques for printed documents are widely used for English language. However, the systems that are implemented to recognize Asian languages struggle to increase the accuracy of recognition. Among other Asian…

计算机视觉与模式识别 · 计算机科学 2014-12-25 G. I. Gunarathna , M. A. P. Chamikara , R. G. Ragel

Inspired by the success of Deep Learning based approaches to English scene text recognition, we pose and benchmark scene text recognition for three Indic scripts - Devanagari, Telugu and Malayalam. Synthetic word images rendered from…

计算机视觉与模式识别 · 计算机科学 2021-04-12 Minesh Mathew , Mohit Jain , CV Jawahar

Chinese text processing systems are using Double Byte Coding , while almost all existing Sanskrit Based Indian Languages have been using Single Byte coding for text processing. Through observation, Chinese Information Processing Technique…

cmp-lg · 计算机科学 2008-02-03 Md Maruf Hasan

Handwritten font generation is important for preserving cultural heritage and creating personalized designs. It adds an authentic and expressive touch to printed materials, making them visually appealing and establishing a stronger…

计算机视觉与模式识别 · 计算机科学 2024-04-05 Preeti P. Bhatt , Jitendra V. Nasriwala , Rakesh R. Savant

Large Language Models (LLMs) have advanced code generation and software automation but remain constrained by inference-time context and lack structured reasoning over code, leaving debugging largely unsolved. While Claude 4.5 Opus achieves…

软件工程 · 计算机科学 2025-12-04 Ishraq Khan , Assad Chowdary , Sharoz Haseeb , Urvish Patel , Yousuf Zaii

Optical character recognition (OCR) is a widely used pattern recognition application in numerous domains. There are several feature-rich, general-purpose OCR solutions available for consumers, which can provide moderate to excellent…

计算机视觉与模式识别 · 计算机科学 2021-05-18 Ayantha Randika , Nilanjan Ray , Xiao Xiao , Allegra Latimer

Optical Character Recognition (OCR) has been a topic of interest for many years. It is defined as the process of digitizing a document image into its constituent characters. Despite decades of intense research, developing OCR with…

计算机视觉与模式识别 · 计算机科学 2017-10-17 Noman Islam , Zeeshan Islam , Nazia Noor

Natural Language Processing (NLP) and especially natural language text analysis have seen great advances in recent times. Usage of deep learning in text processing has revolutionized the techniques for text processing and achieved…

信息检索 · 计算机科学 2020-07-07 Ramchandra Joshi , Purvi Goel , Raviraj Joshi

Recent advances in OCR have shown that an end-to-end (E2E) training pipeline that includes both detection and recognition leads to the best results. However, many existing methods focus primarily on Latin-alphabet languages, often even only…

计算机视觉与模式识别 · 计算机科学 2021-03-31 Jing Huang , Guan Pang , Rama Kovvuri , Mandy Toh , Kevin J Liang , Praveen Krishnan , Xi Yin , Tal Hassner

OCR errors are common in digitised historical archives significantly affecting their usability and value. Generative Language Models (LMs) have shown potential for correcting these errors using the context provided by the corrupted text and…

计算与语言 · 计算机科学 2024-10-01 Jonathan Bourne

Digitization of medical records often relies on smartphone photographs of printed reports, producing images degraded by blur, shadows, and other noise. Conventional OCR systems, optimized for clean scans, perform poorly under such…

信息检索 · 计算机科学 2025-11-18 Nikita Neveditsin , Pawan Lingras , Salil Patil , Swarup Patil , Vijay Mago

After the release of ChatGPT, Large Language Models (LLMs) have gained huge popularity in recent days and thousands of variants of LLMs have been released. However, there is no generative language model for the Nepali language, due to which…

计算与语言 · 计算机科学 2025-06-23 Shushanta Pudasaini , Aman Shakya , Siddhartha Shrestha , Sahil Bhatta , Sunil Thapa , Sushmita Palikhe

Search has for a long time been an important tool for users to retrieve information. Syntactic search is matching documents or objects containing specific keywords like user-history, location, preference etc. to improve the results.…

计算与语言 · 计算机科学 2020-02-26 Arijit Das , Diganta Saha

Translating technical terms into lexically similar, low-resource Indian languages remains a challenge due to limited parallel data and the complexity of linguistic structures. We propose a novel use-case of Sanskrit-based segments for…

计算与语言 · 计算机科学 2026-03-26 Karthika N J , Krishnakant Bhatt , Ganesh Ramakrishnan , Preethi Jyothi

In this paper, we propose a novel method based on character sequence-to-sequence models to correct documents already processed with Optical Character Recognition (OCR) systems. The main contribution of this paper is a set of strategies to…

计算与语言 · 计算机科学 2022-01-26 Juan Ramirez-Orta , Eduardo Xamena , Ana Maguitman , Evangelos Milios , Axel J. Soto

In this paper a fast and novel method is proposed for multi-font multi-size Kannada numeral recognition which is thinning free and without size normalization approach. The different structural feature are used for numeral recognition…

计算机视觉与模式识别 · 计算机科学 2011-11-21 B. V. Dhandra , R. G. Benne , Mallikarjun Hangarge