English
Related papers

Related papers: DohaScript: A Large-Scale Multi-Writer Dataset for…

200 papers

Bangla Sign Language Translation (BdSLT) has been severely constrained so far as the language itself is very low resource. Standard sentence level dataset creation for BdSLT is of immense importance for developing AI based assistive tools…

Computation and Language · Computer Science 2025-11-27 Husne Ara Rubaiyeat , Hasan Mahmud , Md Kamrul Hasan

The recent advances in deep-learning have led to the development of highly sophisticated systems with an unquenchable appetite for data. On the other hand, building good deep-learning models for low-resource languages remains a challenging…

Computation and Language · Computer Science 2024-02-20 Maithili Sabane , Onkar Litake , Aman Chadha

This study investigates the performance of few-shot learning (FSL) approaches in recognizing Bangla handwritten characters and numerals using limited labeled data. It demonstrates the applicability of these methods to scripts with intricate…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Mehedi Ahamed , Radib Bin Kabir , Tawsif Tashwar Dipto , Mueeze Al Mushabbir , Sabbir Ahmed , Md. Hasanul Kabir

We introduce a new dataset for offline Handwritten Text Recognition (HTR) from images of Bangla scripts comprising words, lines, and document-level annotations. The BN-HTRd dataset is based on the BBC Bangla News corpus, meant to act as…

Computer Vision and Pattern Recognition · Computer Science 2022-06-22 Md. Ataur Rahman , Nazifa Tabassum , Mitu Paul , Riya Pal , Mohammad Khairul Islam

How can an end-user provide feedback if a deployed structured prediction model generates inconsistent output, ignoring the structural complexity of human language? This is an emerging topic with recent progress in synthetic or constrained…

Artificial Intelligence · Computer Science 2021-12-17 Niket Tandon , Aman Madaan , Peter Clark , Keisuke Sakaguchi , Yiming Yang

Handwritten character classification in the Bengali script is a significant challenge due to the complexity and variability of the characters. The models commonly used for classification are often computationally expensive and data-hungry,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-04 Rafi Hassan Chowdhury , Naimul Haque , Kaniz Fatiha

This study developed a new Bangla abstractive summarization dataset to generate concise summaries of Bangla articles from diverse sources. Most existing studies in this field have concentrated on news articles, where journalists usually…

Computation and Language · Computer Science 2025-12-17 Md. Tanzim Ferdous , Naeem Ahsan Chowdhury , Prithwiraj Bhattacharjee

Statistical watermarking is a common approach for verifying whether text was written by a language model. Most existing schemes assume autoregressive generation, where tokens are produced left to right and contextual hashing is well…

Computation and Language · Computer Science 2026-05-08 Mohd Ruhul Ameen , Akif Islam , Nadim Mahmud , Md. Ekramul Hamid

The application of handwritten text recognition to historical works is highly dependant on accurate text line retrieval. A number of systems utilizing a robust baseline detection paradigm have emerged recently but the advancement of layout…

Computer Vision and Pattern Recognition · Computer Science 2019-07-10 Benjamin Kiessling , Daniel Stökl Ben Ezra , Matthew Thomas Miller

Despite considerable progress in handwritten text recognition, paragraph-level handwritten text recognition, especially in low-resource languages, such as Hindi, Urdu and similar scripts, remains a challenging problem. These languages,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Sayantan Dey , Alireza Alaei , Partha Pratim Roy

Despite its significance, Arabic, a linguistically rich and morphologically complex language, faces the challenge of being under-resourced. The scarcity of large annotated datasets hampers the development of accurate tools for subjectivity…

Computation and Language · Computer Science 2026-03-02 Slimane Bellaouar , Attia Nehar , Soumia Souffi , Mounia Bouameur

Current Text-to-Speech models pose a multilingual challenge, where most of the models traditionally focus on English and European languages, thereby hurting the potential to provide access to information to many more people. To address this…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-21 Jaskaran Singh , Amartya Roy Chowdhury , Raghav Prabhakar , Varshul C. W

India has a rich linguistic landscape with languages from 4 major language families spoken by over a billion people. 22 of these languages are listed in the Constitution of India (referred to as scheduled languages) are the focus of this…

Recognition of handwritten Bangla compound characters remains a challenging problem due to complex character structures, large intra-class variation, and limited availability of high-quality annotated data. Existing Bangla handwritten…

Computer Vision and Pattern Recognition · Computer Science 2026-05-13 Md. Sultan Al Rayhan

Segmentation of Arabic manuscripts into lines of text and words is an important step to make recognition systems more efficient and accurate. The problem of segmentation into text lines is solved since there are carefully annotated dataset…

Computation and Language · Computer Science 2023-12-14 Hakim Bouchal , Ahror Belaid

The Indus script is one of the major undeciphered scripts of the ancient world. The small size of the corpus, the absence of bilingual texts, and the lack of definite knowledge of the underlying language has frustrated efforts at…

Computation and Language · Computer Science 2015-05-13 Nisha Yadav , Hrishikesh Joglekar , Rajesh P. N. Rao , M. N. Vahia , Iravatham Mahadevan , R. Adhikari

Full-duplex spoken dialogue systems can model natural conversational behaviours such as interruptions, overlaps, and backchannels, yet such systems remain largely unexplored for Indian languages. We present the first open, reproducible…

Computation and Language · Computer Science 2026-05-26 Bhaskar Singh , Shobhit Banga , Mahima Manik , Pranav Sharma

Handwritten Text Recognition (HTR) under limited labeled data remains a challenging problem, particularly for Arabic-script languages. Although modern sequence-based recognizers perform well in high-resource settings, their accuracy…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Sana Al-azzawi , Elisa Barney , Marcus Liwicki

In this paper, we introduce the first and largest Hindi text corpus, named BHAAV, which means emotions in Hindi, for analyzing emotions that a writer expresses through his characters in a story, as perceived by a narrator/reader. The corpus…

Computation and Language · Computer Science 2019-10-10 Yaman Kumar , Debanjan Mahata , Sagar Aggarwal , Anmol Chugh , Rajat Maheshwari , Rajiv Ratn Shah

Diacritics are orthographic marks that clarify pronunciation, distinguish similar words, or alter meaning. They play a central role in many writing systems, yet their impact on language technology has not been systematically quantified…

Computation and Language · Computer Science 2026-03-31 Adi Cohen , Yuval Pinter
‹ Prev 1 3 4 5 6 7 10 Next ›