English
Related papers

Related papers: ADI-20: Arabic Dialect Identification dataset and …

200 papers

This paper proposes a Dialect Identification (DID) approach inspired by the Connectionist Temporal Classification (CTC) loss function as used in Automatic Speech Recognition (ASR). CTC-DID frames the dialect identification task as a…

Computation and Language · Computer Science 2026-01-21 Muhammad Umar Farooq , Oscar Saz

We investigate different approaches for dialect identification in Arabic broadcast speech, using phonetic, lexical features obtained from a speech recognition system, and acoustic features using the i-vector framework. We studied both…

Computation and Language · Computer Science 2016-08-12 Ahmed Ali , Najim Dehak , Patrick Cardinal , Sameer Khurana , Sree Harsha Yella , James Glass , Peter Bell , Steve Renals

The Arabic language suffers from a great shortage of datasets suitable for training deep learning models, and the existing ones include general non-specialized classifications. In this work, we introduce a new Arab medical dataset, which…

Computation and Language · Computer Science 2021-07-06 Jaafar Hammoud , Aleksandra Vatian , Natalia Dobrenko , Nikolai Vedernikov , Anatoly Shalyto , Natalia Gusarova

This work consists of creating a system of the Computer Assisted Language Learning (CALL) based on a system of Automatic Speech Recognition (ASR) for the Arabic language using the tool CMU Sphinx3 [1], based on the approach of HMM. To this…

Computation and Language · Computer Science 2012-05-16 Naim Terbeh , Mounir Zrigui

This study presents systems submitted by the University of Texas at Dallas, Center for Robust Speech Systems (UTD-CRSS) to the MGB-3 Arabic Dialect Identification (ADI) subtask. This task is defined to discriminate between five dialects of…

Audio and Speech Processing · Electrical Eng. & Systems 2017-10-03 Ahmet E. Bulut , Qian Zhang , Chunlei Zhang , Fahimeh Bahmaninezhad , John H. L. Hansen

Pre-trained Language Models (PLMs) are integral to many modern natural language processing (NLP) systems. Although multilingual models cover a wide range of languages, they often grapple with challenges like high inference costs and a lack…

Computation and Language · Computer Science 2024-07-19 Murtadha Ahmed , Saghir Alfasly , Bo Wen , Jamaal Qasem , Mohammed Ahmed , Yunfeng Liu

We present improvements in automatic speech recognition (ASR) for Somali, a currently extremely under-resourced language. This forms part of a continuing United Nations (UN) effort to employ ASR-based keyword spotting systems to support…

Computation and Language · Computer Science 2019-07-09 Astik Biswas , Raghav Menon , Ewald van der Westhuizen , Thomas Niesler

The Arabic language is characterized by a rich tapestry of regional dialects that differ substantially in phonetics and lexicon, reflecting the geographic and cultural diversity of its speakers. Despite the availability of many…

The paper introduces and publicly releases (Data download link available after acceptance) CAFE -- the first Code-switching dataset between Algerian dialect, French, and english languages. The CAFE speech data is unique for (a) its…

Arabic dialect recognition presents a significant challenge in speech technology due to the linguistic diversity of Arabic and the scarcity of large annotated datasets, particularly for underrepresented dialects. This research investigates…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-27 Ghazal Al-Shwayyat , Omer Nezih Gerek

We present a machine learning approach that ranked on the first place in the Arabic Dialect Identification (ADI) Closed Shared Tasks of the 2018 VarDial Evaluation Campaign. The proposed approach combines several kernels using multiple…

Computation and Language · Computer Science 2018-07-31 Andrei M. Butnaru , Radu Tudor Ionescu

This paper introduces Mixat: a dataset of Emirati speech code-mixed with English. Mixat was developed to address the shortcomings of current speech recognition resources when applied to Emirati speech, and in particular, to bilignual…

Computation and Language · Computer Science 2024-05-07 Maryam Al Ali , Hanan Aldarmaki

Arabic dialect identification is a specific task of natural language processing, aiming to automatically predict the Arabic dialect of a given text. Arabic dialect identification is the first step in various natural language processing…

Computation and Language · Computer Science 2020-09-29 Maha J. Althobaiti

The focus of language model evaluation has transitioned towards reasoning and knowledge-intensive tasks, driven by advancements in pretraining large models. While state-of-the-art models are partially trained on large Arabic texts,…

Arabic, with its rich diversity of dialects, remains significantly underrepresented in Large Language Models, particularly in dialectal variations. We address this gap by introducing seven synthetic datasets in dialects alongside Modern…

Computation and Language · Computer Science 2024-12-19 Basel Mousi , Nadir Durrani , Fatema Ahmad , Md. Arid Hasan , Maram Hasanain , Tameem Kabbani , Fahim Dalvi , Shammur Absar Chowdhury , Firoj Alam

Natural Language Processing (NLP) is today a very active field of research and innovation. Many applications need however big sets of data for supervised learning, suitably labelled for the training purpose. This includes applications for…

Computation and Language · Computer Science 2021-02-23 ElMehdi Boujou , Hamza Chataoui , Abdellah El Mekki , Saad Benjelloun , Ikram Chairi , Ismail Berrada

This paper addresses critical gaps in Arabic language model evaluation by establishing comprehensive theoretical guidelines and introducing a novel evaluation framework. We first analyze existing Arabic evaluation datasets, identifying…

Computation and Language · Computer Science 2025-06-03 Serry Sibaee , Omer Nacar , Adel Ammar , Yasser Al-Habashi , Abdulrahman Al-Batati , Wadii Boulila

We describe findings of the third Nuanced Arabic Dialect Identification Shared Task (NADI 2022). NADI aims at advancing state of the art Arabic NLP, including on Arabic dialects. It does so by affording diverse datasets and modeling…

Computation and Language · Computer Science 2022-10-24 Muhammad Abdul-Mageed , Chiyu Zhang , AbdelRahim Elmadany , Houda Bouamor , Nizar Habash

We describe the winning system for Task 2 of the KSAA-2026 Shared Task on Arabic Speech Dictation with Automatic Diacritization. The task requires producing fully diacritized Arabic text from speech audio and undiacritized transcripts, with…

Computation and Language · Computer Science 2026-05-26 Meshal Alamr , Hassan Alqaeri , Abdullah Aldahlawi

This work is an attempt to introduce a comprehensive benchmark for Arabic speech recognition, specifically tailored to address the challenges of telephone conversations in Arabic language. Arabic, characterized by its rich dialectal…

Artificial Intelligence · Computer Science 2024-05-31 Qusai Abo Obaidah , Muhy Eddin Za'ter , Adnan Jaljuli , Ali Mahboub , Asma Hakouz , Bashar Al-Rfooh , Yazan Estaitia