English
Related papers

Related papers: A Benchmark and Dataset for Post-OCR text correcti…

200 papers

SiDiaC, the first comprehensive Sinhala Diachronic Corpus, covers a historical span from the 5th to the 20th century CE. SiDiaC comprises 58k words across 46 literary works, annotated carefully based on the written date, after filtering…

Computation and Language · Computer Science 2026-05-19 Nevidu Jayatilleke , Nisansa de Silva

Automatic speech recognition (ASR) performs well for high-resource languages with abundant paired audio-transcript data, but its accuracy degrades sharply for most languages due to limited publicly available aligned data. To this end, we…

Computation and Language · Computer Science 2026-05-12 Antonis Asonitis , Luca A. Lanzendörfer , Frédéric Berdoz , Roger Wattenhofer

Segmentation of a text-document into lines, words and characters, which is considered to be the crucial pre-processing stage in Optical Character Recognition (OCR) is traditionally carried out on uncompressed documents, although most of the…

Computer Vision and Pattern Recognition · Computer Science 2014-04-01 Mohammed Javed , P. Nagabhushan , B. B. Chaudhuri

Source Code Summarization is the task of writing short, natural language descriptions of source code. The main use for these descriptions is in software documentation e.g. the one-sentence Java method descriptions in JavaDocs. Code…

Computation and Language · Computer Science 2019-04-05 Alexander LeClair , Collin McMillan

This paper presents first benchmark corpus of Sanskrit Pratyaya (suffix) and inflectional words (padas) formed due to suffixes along with neural network based approaches to process the formation and splitting of inflectional words.…

Computation and Language · Computer Science 2024-09-05 Arun Kumar Singh , Sushant Dave , Prathosh A. P. , Brejesh Lall , Shresth Mehta

Scene text recognition in low-resource Indian languages is challenging because of complexities like multiple scripts, fonts, text size, and orientations. In this work, we investigate the power of transfer learning for all the layers of deep…

Computer Vision and Pattern Recognition · Computer Science 2022-01-11 Sanjana Gunna , Rohit Saluja , C. V. Jawahar

This paper is devoted to the adaptation of generative large language models for the Tajik language, a low-resource language with Cyrillic script. To overcome the shortage of digital text resources, the author created and publicly released…

Computation and Language · Computer Science 2026-05-06 Mullosharaf K. Arabov

This paper describes neural network based approaches to the process of the formation and splitting of word-compounding, respectively known as the Sandhi and Vichchhed, in Sanskrit language. Sandhi is an important idea essential to…

Computation and Language · Computer Science 2024-09-05 Sushant Dave , Arun Kumar Singh , Prathosh A. P. , Brejesh Lall

Manchu, a critically endangered language essential for understanding early modern Eastern Eurasian history, lacks effective OCR systems that can handle real-world historical documents. This study develops high-performing OCR systems by…

Computer Vision and Pattern Recognition · Computer Science 2025-07-10 Yan Hon Michael Chung , Donghyeok Choi

Industrial Retrieval-Augmented Generation (RAG) systems depend on optical character recognition (OCR) to transform visual documents into text. Existing OCR benchmarks rely on character-level metrics, which inadequately measure downstream…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Lin Sun , Wang Dexian , Jingang Huang , Linglin Zhang , Change Jia , Zhengwei Cheng , Xiangzheng Zhang

Poetry generation in Sanskrit typically requires the verse to be semantically coherent and adhere to strict prosodic rules. In Sanskrit prosody, every line of a verse is typically a fixed length sequence of syllables adhering to prescribed…

Computation and Language · Computer Science 2026-03-26 Manoj Balaji Jagadeeshan , Atul Singh , Nallani Chakravartula Sahith , Amrith Krishna , Pawan Goyal

Spelling error correction is the task of identifying and rectifying misspelled words in texts. It is a potential and active research topic in Natural Language Processing because of numerous applications in human language understanding. The…

Computation and Language · Computer Science 2024-12-03 Mehedi Hasan Bijoy , Nahid Hossain , Salekul Islam , Swakkhar Shatabda

We introduce UniversalCEFR, a large-scale multilingual and multidimensional dataset of texts annotated with CEFR (Common European Framework of Reference) levels in 13 languages. To enable open research in automated readability and language…

Kazakh is a Turkic language using the Arabic, Cyrillic, and Latin scripts, making it unique in terms of optical character recognition (OCR). Work on OCR for low-resource Kazakh scripts is very scarce, and no OCR benchmarks or images exist…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Henry Gagnier , Sophie Gagnier , Ashwin Kirubakaran

Workplace meetings are vital to organizational collaboration, yet a large percentage of meetings are rated as ineffective. To help improve meeting effectiveness by understanding if the conversation is on topic, we create a comprehensive…

Computation and Language · Computer Science 2024-11-05 Yaran Fan , Jamie Pool , Senja Filipi , Ross Cutler

Indonesia is rich in languages and scripts. However, most NLP progress has been made using romanized text. In this paper, we present NusaAksara, a novel public benchmark for Indonesian languages that includes their original scripts. Our…

Computation and Language · Computer Science 2025-08-06 Muhammad Farid Adilazuarda , Musa Izzanardi Wijanarko , Lucky Susanto , Khumaisa Nur'aini , Derry Wijaya , Alham Fikri Aji

Semantic code search is the task of retrieving relevant code given a natural language query. While related to other information retrieval tasks, it requires bridging the gap between the language used in code (often abbreviated and highly…

Machine Learning · Computer Science 2020-06-09 Hamel Husain , Ho-Hsiang Wu , Tiferet Gazit , Miltiadis Allamanis , Marc Brockschmidt

Developing benchmark datasets for low-resource languages poses significant challenges, primarily due to the limited availability of native linguistic experts and the substantial time and cost involved in annotation. Given these challenges,…

Computation and Language · Computer Science 2025-10-28 Rahul Ranjan , Mahendra Kumar Gurve , Anuj , Nitin , Yamuna Prasad

This paper introduces a novel approach to post-Optical Character Recognition Correction (POC) for handwritten Cyrillic text, addressing a significant gap in current research methodologies. This gap is due to the lack of large text corporas…

Computer Vision and Pattern Recognition · Computer Science 2023-11-28 Evgenii Davydkin , Aleksandr Markelov , Egor Iuldashev , Anton Dudkin , Ivan Krivorotov

Diacritization of Arabic text is both an interesting and a challenging problem at the same time with various applications ranging from speech synthesis to helping students learning the Arabic language. Like many other tasks or problems in…

Computation and Language · Computer Science 2019-05-07 Ali Fadel , Ibraheem Tuffaha , Bara' Al-Jawarneh , Mahmoud Al-Ayyoub
‹ Prev 1 8 9 10 Next ›