English
Related papers

Related papers: Comprehensive Benchmark Datasets for Amharic Scene…

200 papers

In spite of the recent progress in speech processing, the majority of world languages and dialects remain uncovered. This situation only furthers an already wide technological divide, thereby hindering technological and socioeconomic…

This research paper presents a part-of-speech (POS) annotated dataset and tagger tool for the low-resource Uzbek language. The dataset includes 12 tags, which were used to develop a rule-based POS-tagger tool. The corpus text used in the…

Computation and Language · Computer Science 2023-03-02 Maksud Sharipov , Elmurod Kuriyozov , Ollabergan Yuldashev , Ogabek Sobirov

Yor\`ub\'a is a widely spoken West African language with a writing system rich in tonal and orthographic diacritics. With very few exceptions, diacritics are omitted from electronic texts, due to limited device and application support.…

Computation and Language · Computer Science 2018-10-31 Iroro Orife

Hausa language belongs to the Afroasiatic phylum, and with more first-language speakers than any other sub-Saharan African language. With a majority of its speakers residing in the Northern and Southern areas of Nigeria and the Republic of…

Computation and Language · Computer Science 2021-02-18 Isa Inuwa-Dutse

We present ADI-20, an extension of the previously published ADI-17 Arabic Dialect Identification (ADI) dataset. ADI-20 covers all Arabic-speaking countries' dialects. It comprises 3,556 hours from 19 Arabic dialects in addition to Modern…

Computation and Language · Computer Science 2025-11-14 Haroun Elleuch , Salima Mdhaffar , Yannick Estève , Fethi Bougares

This study thoroughly investigates how well deep learning models can recognize Arabic handwritten text for person biometric identification. It compares three advanced architectures -- ResNet50, MobileNetV2, and EfficientNetB7 -- using three…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Mazen Balat , Youssef Mohamed , Ahmed Heakl , Ahmed Zaky

The Houma Alliance Book, one of history's earliest calligraphic examples, was unearthed in the 1970s. These artifacts were meticulously organized, reproduced, and copied by the Shanxi Provincial Institute of Cultural Relics. However,…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Xiaoyu Yuan , Xiaohua Huang , Zibo Zhang , Yabo Sun

With the rise of generative text-to-speech models, distinguishing between real and synthetic speech has become challenging, especially for Arabic that have received limited research attention. Most spoof detection efforts have focused on…

Computation and Language · Computer Science 2025-09-30 Mohamed Maged , Alhassan Ehab , Ali Mekky , Besher Hassan , Shady Shehata

The development of Urdu scene text detection, recognition, and Visual Question Answering (VQA) technologies is crucial for advancing accessibility, information retrieval, and linguistic diversity in digital content, facilitating better…

Computer Vision and Pattern Recognition · Computer Science 2024-05-22 Hiba Maryam , Ling Fu , Jiajun Song , Tajrian ABM Shafayet , Qidi Luo , Xiang Bai , Yuliang Liu

The recognition of cursive script is regarded as a subtle task in optical character recognition due to its varied representation. Every cursive script has different nature and associated challenges. As Urdu is one of cursive language that…

Computer Vision and Pattern Recognition · Computer Science 2017-05-17 Saad Bin Ahmed , Saeeda Naz , Salahuddin Swati , Muhammad Imran Razzak

In this work, we present AfriHuBERT, an extension of mHuBERT-147, a compact self-supervised learning (SSL) model pretrained on 147 languages. While mHuBERT-147 covered 16 African languages, we expand this to 1,226 through continued…

Computation and Language · Computer Science 2025-06-03 Jesujoba O. Alabi , Xuechen Liu , Dietrich Klakow , Junichi Yamagishi

The proliferation of online offensive language necessitates the development of effective detection mechanisms, especially in multilingual contexts. This study addresses the challenge by developing and introducing novel datasets for…

Computation and Language · Computer Science 2024-06-07 Saminu Mohammad Aliyu , Gregory Maksha Wajiga , Muhammad Murtala

Sign language recognition is a challenging problem where signs are identified by simultaneous local and global articulations of multiple sources, i.e. hand shape and orientation, hand movements, body posture, and facial expressions. Solving…

Computer Vision and Pattern Recognition · Computer Science 2020-10-20 Ozge Mercanoglu Sincan , Hacer Yalim Keles

We introduce ArVoice, a multi-speaker Modern Standard Arabic (MSA) speech corpus with diacritized transcriptions, intended for multi-speaker speech synthesis, and can be useful for other tasks such as speech-based diacritic restoration,…

Computation and Language · Computer Science 2025-05-28 Hawau Olamide Toyin , Rufael Marew , Humaid Alblooshi , Samar M. Magdy , Hanan Aldarmaki

In daily communications, Arabs use local dialects which are hard to identify automatically using conventional classification methods. The dialect identification challenging task becomes more complicated when dealing with an under-resourced…

Computation and Language · Computer Science 2017-03-30 Soumia Bougrine , Hadda Cherroun , Djelloul Ziadi

We present the first Africentric SemEval Shared task, Sentiment Analysis for African Languages (AfriSenti-SemEval) - The dataset is available at https://github.com/afrisenti-semeval/afrisent-semeval-2023. AfriSenti-SemEval is a sentiment…

Current Machine Translation (MT) systems for Arabic often struggle to account for dialectal diversity, frequently homogenizing dialectal inputs into Modern Standard Arabic (MSA) and offering limited user control over the target vernacular.…

Computation and Language · Computer Science 2026-04-09 Afroza Nowshin , Prithweeraj Acharjee Porag , Haziq Jeelani , Fayeq Jeelani Syed

This paper describes the acquisition, preprocessing, segmentation, and alignment of an Amharic-English parallel corpus. It will be helpful for machine translation of a low-resource language, Amharic. We freely released the corpus for…

Computation and Language · Computer Science 2022-05-04 Andargachew Mekonnen Gezmu , Andreas Nürnberger , Tesfaye Bayu Bati

Recent advances in Deep Learning and Computer Vision have been successfully leveraged to serve marginalized communities in various contexts. One such area is Sign Language - a primary means of communication for the deaf community. However,…

Computer Vision and Pattern Recognition · Computer Science 2024-01-23 Haz Sameen Shahgir , Khondker Salman Sayeed , Md Toki Tahmid , Tanjeem Azwad Zaman , Md. Zarif Ul Alam

Large Language Models (LLMs) are transforming Natural Language Processing (NLP), but their benefits are largely absent for Africa's 2,000 low-resource languages. This paper comparatively analyzes African language coverage across six LLMs,…

‹ Prev 1 4 5 6 7 8 10 Next ›