English
Related papers

Related papers: ARCADE: A City-Scale Corpus for Fine-Grained Arabi…

200 papers

Arabic is a complex language with many varieties and dialects spoken by over 450 millions all around the world. Due to the linguistic diversity and variations, it is challenging to build a robust and generalized ASR system for Arabic. In…

Computation and Language · Computer Science 2023-10-30 Abdul Waheed , Bashar Talafha , Peter Sullivan , AbdelRahim Elmadany , Muhammad Abdul-Mageed

On annotating multi-dialect Arabic datasets, it is common to randomly assign the samples across a pool of native Arabic speakers. Recent analyses recommended routing dialectal samples to native speakers of their respective dialects to build…

Computation and Language · Computer Science 2024-06-10 Amr Keleg , Walid Magdy , Sharon Goldwater

Arabic is one of the oldest languages still in use today. As a result, several Arabic-speaking regions have developed dialects that are unique to them. Dialect and emotion recognition have various uses in Arabic text analysis, such as…

Computation and Language · Computer Science 2025-02-14 Nasser A Alsadhan

We present QADI, an automatically collected dataset of tweets belonging to a wide range of country-level Arabic dialects -covering 18 different countries in the Middle East and North Africa region. Our method for building this dataset…

Computation and Language · Computer Science 2020-05-18 Ahmed Abdelali , Hamdy Mubarak , Younes Samih , Sabit Hassan , Kareem Darwish

We present ZAEBUC-Spoken, a multilingual multidialectal Arabic-English speech corpus. The corpus comprises twelve hours of Zoom meetings involving multiple speakers role-playing a work situation where Students brainstorm ideas for a certain…

Computation and Language · Computer Science 2024-03-28 Injy Hamed , Fadhl Eryani , David Palfreyman , Nizar Habash

The automatic classification of Arabic dialects is an ongoing research challenge, which has been explored in recent work that defines dialects based on increasingly limited geographic areas like cities and provinces. This paper focuses on a…

Computation and Language · Computer Science 2021-10-04 Abdulkareem Alsudais , Wafa Alotaibi , Faye Alomary

This paper presents the annotation guidelines of the Balanced Arabic Readability Evaluation Corpus (BAREC), a large-scale resource for fine-grained sentence-level readability assessment in Arabic. BAREC includes 69,441 sentences (1M+ words)…

Computation and Language · Computer Science 2025-06-12 Nizar Habash , Hanada Taha-Thomure , Khalid N. Elmadani , Zeina Zeino , Abdallah Abushmaes

Modern Arabic ASR systems such as wav2vec 2.0 excel at word- and sentence-level transcription, yet struggle to classify isolated letters. In this study, we show that this phoneme-level task, crucial for language learning, speech therapy,…

Computation and Language · Computer Science 2025-08-28 Hadi Zaatiti , Hatem Hajri , Osama Abdullah , Nader Masmoudi

Arabic is one of the most important and growing languages in the world. With the rise of social media platforms such as Twitter, Arabic spoken dialects have become more in use. In this paper, we describe our approach on the NADI Shared Task…

Computation and Language · Computer Science 2020-11-16 Ahmad Beltagy , Abdelrahman Wael , Omar ElSherief

Cross-Lingual SynthDocs is a large-scale synthetic corpus designed to address the scarcity of Arabic resources for Optical Character Recognition (OCR) and Document Understanding (DU). The dataset comprises over 2.5 million of samples,…

Computation and Language · Computer Science 2025-11-10 Haneen Al-Homoud , Asma Ibrahim , Murtadha Al-Jubran , Fahad Al-Otaibi , Yazeed Al-Harbi , Daulet Toibazar , Kesen Wang , Pedro J. Moreno

Identifying hate speech content in the Arabic language is challenging due to the rich quality of dialectal variations. This study introduces a multilabel hate speech dataset in the Arabic language. We have collected 10000 Arabic tweets and…

Computation and Language · Computer Science 2025-05-26 Wajdi Zaghouani , Md. Rafiul Biswas

Automatic Arabic Dialect Identification (ADI) of text has gained great popularity since it was introduced in the early 2010s. Multiple datasets were developed, and yearly shared tasks have been running since 2018. However, ADI systems are…

Computation and Language · Computer Science 2023-10-23 Amr Keleg , Walid Magdy

We present ASCAT (Arabic Scientific Corpus for Advanced Translation), a high-quality English-Arabic parallel benchmark corpus designed for scientific translation evaluation constructed through a systematic multi-engine translation and human…

Computation and Language · Computer Science 2026-04-02 Serry Sibaee , Khloud Al Jallad , Zineb Yousfi , Israa Elsayed Elhosiny , Yousra El-Ghawi , Batool Balah , Omer Nacar

Developing robust automatic speech recognition (ASR) systems for Arabic requires effective strategies to manage its diversity. Existing ASR systems mainly cover the modern standard Arabic (MSA) variety and few high-resource dialects, but…

Computation and Language · Computer Science 2025-06-02 Amirbek Djanibekov , Hawau Olamide Toyin , Raghad Alshalan , Abdullah Alitr , Hanan Aldarmaki

This article presents morphologically-annotated Yemeni, Sudanese, Iraqi, and Libyan Arabic dialects Lisan corpora. Lisan features around 1.2 million tokens. We collected the content of the corpora from several social media platforms. The…

Computation and Language · Computer Science 2022-12-20 Mustafa Jarrar , Fadi A Zaraket , Tymaa Hammouda , Daanish Masood Alavi , Martin Waahlisch

In order to successfully annotate the Arabic speech con- tent found in open-domain media broadcasts, it is essential to be able to process a diverse set of Arabic dialects. For the 2017 Multi-Genre Broadcast challenge (MGB-3) there were two…

Computation and Language · Computer Science 2017-09-04 Suwon Shon , Ahmed Ali , James Glass

Semantic segmentation is a core component of discourse analysis, yet existing models are primarily developed and evaluated on high-resource written text, limiting their effectiveness on low-resource spoken varieties. In particular,…

Computation and Language · Computer Science 2026-05-08 Kirill Chirkunov , Younes Samih , Abed Alhakim Freihat , Hanan Aldarmaki

In daily communications, Arabs use local dialects which are hard to identify automatically using conventional classification methods. The dialect identification challenging task becomes more complicated when dealing with an under-resourced…

Computation and Language · Computer Science 2017-03-30 Soumia Bougrine , Hadda Cherroun , Djelloul Ziadi

We present our effort to create a large Multi-Layered representational repository of Linguistic Code-Switched Arabic data. The process involves developing clear annotation standards and Guidelines, streamlining the annotation process, and…

Computation and Language · Computer Science 2019-10-01 Mona Diab , Mahmoud Ghoneim , Abdelati Hawwari , Fahad AlGhamdi , Nada AlMarwani , Mohamed Al-Badrashiny

With the rise of generative text-to-speech models, distinguishing between real and synthetic speech has become challenging, especially for Arabic that have received limited research attention. Most spoof detection efforts have focused on…

Computation and Language · Computer Science 2025-09-30 Mohamed Maged , Alhassan Ehab , Ali Mekky , Besher Hassan , Shady Shehata