English
Related papers

Related papers: Comprehensive Benchmark Datasets for Amharic Scene…

200 papers

The objective of the paper is to recognize handwritten samples of basic Bangla characters using Tesseract open source Optical Character Recognition (OCR) engine under Apache License 2.0. Handwritten data samples containing isolated Bangla…

Computer Vision and Pattern Recognition · Computer Science 2010-03-31 Sandip Rakshit , Debkumar Ghosal , Tanmoy Das , Subhrajit Dutta , Subhadip Basu

The ultimate aim of handwriting recognition is to make computers able to read and/or authenticate human written texts, with a performance comparable to or even better than that of humans. Reading means that the computer is given a piece of…

Computer Vision and Pattern Recognition · Computer Science 2012-06-26 Manal A. Abdullah , Lulwah M. Al-Harigy , Hanadi H. Al-Fraidi

Sign language is the primary approach of communication for the Deaf and Hard-of-Hearing (DHH) community. While there are numerous benchmarks for high-resource sign languages, low-resource languages like Arabic remain underrepresented.…

Computation and Language · Computer Science 2026-04-01 Mohammad Amer Khalil , Raghad Nahas , Ahmad Nassar , Khloud Al Jallad

Video-to-text and text-to-video retrieval are dominated by English benchmarks (e.g. DiDeMo, MSR-VTT) and recent multilingual corpora (e.g. RUDDER), yet Arabic remains underserved, lacking localized evaluation metrics. We introduce a…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Mohamed Eltahir , Osamah Sarraj , Abdulrahman Alfrihidi , Taha Alshatiri , Mohammed Khurd , Mohammed Bremoo , Tanveer Hussain

The advancement of speech technologies has been remarkable, yet its integration with African languages remains limited due to the scarcity of African speech corpora. To address this issue, we present AfroDigits, a minimalist,…

Determination of mispronunciations and ensuring feedback to users are maintained by computer-assisted language learning (CALL) systems. In this work, we introduce an ensemble model that defines the mispronunciation of Arabic phonemes and…

Sound · Computer Science 2023-01-05 Sukru Selim Calik , Ayhan Kucukmanisa , Zeynep Hilal Kilimci

Despite the importance of handwritten numeral classification, a robust and effective method for a widely used language like Arabic is still due. This study focuses to overcome two major limitations of existing works: data diversity and…

Computer Vision and Pattern Recognition · Computer Science 2019-08-07 S. M. A. Sharif , Ghulam Mujtaba , S. M. Nadim Uddin

The largest dataset of Arabic speech mispronunciation detections in Egyptian dialogues is introduced. The dataset is composed of annotated audio files representing the top 100 words that are most frequently used in the Arabic language,…

Computation and Language · Computer Science 2021-11-03 Salah A. Aly , Abdelrahman Salah , Hesham M. Eraqi

We present a novel Automatic Speech Recognition (ASR) dataset for the Oromo language, a widely spoken language in Ethiopia and neighboring regions. The dataset was collected through a crowd-sourcing initiative, encompassing a diverse range…

Computation and Language · Computer Science 2025-02-04 Turi Abu , Ying Shi , Thomas Fang Zheng , Dong Wang

Africa is home to over 2,000 languages from more than six language families and has the highest linguistic diversity among all continents. These include 75 languages with at least one million speakers each. Yet, there is little NLP research…

Arabic Rhetoric is the field of Arabic linguistics which governs the art and science of conveying a message with greater beauty, impact and persuasiveness. The field is as ancient as the Arabic language itself and is found extensively in…

Computation and Language · Computer Science 2025-07-30 Mandar Marathe

Automatic Arabic handwritten recognition is one of the recently studied problems in the field of Machine Learning. Unlike Latin languages, Arabic is a Semitic language that forms a harder challenge, especially with variability of patterns…

Computer Vision and Pattern Recognition · Computer Science 2022-11-07 Mais Alheraki , Rawan Al-Matham , Hend Al-Khalifa

A prototype system for the transliteration of diacritics-less Arabic manuscripts at the sub-word or part of Arabic word (PAW) level is developed. The system is able to read sub-words of the input manuscript using a set of skeleton-based…

Computer Vision and Pattern Recognition · Computer Science 2013-06-27 Reza Farrahi Moghaddam , Mohamed Cheriet , Thomas Milo , Robert Wisnovsky

The performance of Artificial Intelligence (AI) systems fundamentally depends on high-quality training data. However, low-resource languages like Arabic suffer from severe data scarcity. Moreover, the absence of child-specific speech…

Computation and Language · Computer Science 2025-10-28 Mouhand Alkadri , Dania Desouki , Khloud Al Jallad

This paper presents our approach to multi-label emotion detection in Hausa, a low-resource African language, for SemEval Track A. We fine-tuned AfriBERTa, a transformer-based model pre-trained on African languages, to classify Hausa text…

Computation and Language · Computer Science 2025-06-24 Sani Abdullahi Sani , Salim Abubakar , Falalu Ibrahim Lawan , Abdulhamid Abubakar , Maryam Bala

In this paper, we describe a spoken Arabic dialect identification (ADI) model for Arabic that consistently outperforms previously published results on two benchmark datasets: ADI-5 and ADI-17. We explore two architectural variations: ResNet…

Computation and Language · Computer Science 2023-10-24 Ajinkya Kulkarni , Hanan Aldarmaki

Developing effective scene text detection and recognition models hinges on extensive training data, which can be both laborious and costly to obtain, especially for low-resourced languages. Conventional methods tailored for Latin characters…

Computer Vision and Pattern Recognition · Computer Science 2024-10-25 Vannkinh Nom , Souhail Bakkali , Muhammad Muzzamil Luqman , Mickaël Coustaty , Jean-Marc Ogier

Word embeddings are a core component of modern natural language processing systems, making the ability to thoroughly evaluate them a vital task. We describe DiaLex, a benchmark for intrinsic evaluation of dialectal Arabic word embedding.…

Automatic Arabic diacritization is useful in many applications, ranging from reading support for language learners to accurate pronunciation predictor for downstream tasks like speech synthesis. While most of the previous works focused on…

Computation and Language · Computer Science 2023-08-01 Parnia Bahar , Mattia Di Gangi , Nick Rossenbach , Mohammad Zeineldeen

Automated Essay Scoring (AES) holds significant promise in the field of education, helping educators to mark larger volumes of essays and provide timely feedback. However, Arabic AES research has been limited by the lack of publicly…

Computation and Language · Computer Science 2024-07-17 Rayed Ghazawi , Edwin Simpson
‹ Prev 1 3 4 5 6 7 10 Next ›