English
Related papers

Related papers: ARCADE: A City-Scale Corpus for Fine-Grained Arabi…

200 papers

This paper addresses the challenge of learning to recite the Quran for non-Arabic speakers. We explore the possibility of crowdsourcing a carefully annotated Quranic dataset, on top of which AI models can be built to simplify the learning…

Sound · Computer Science 2024-05-07 Raghad Salameh , Mohamad Al Mdfaa , Nursultan Askarbekuly , Manuel Mazzara

Mental health disorders affect millions worldwide, yet early detection remains a major challenge, particularly for Arabic-speaking populations where resources are limited and mental health discourse is often discouraged due to cultural…

Computation and Language · Computer Science 2025-11-06 Saad Mankarious , Ayah Zirikly

Despite growing interest in Quranic data research, existing Quran datasets remain limited in both scale and diversity. To address this gap, we present Tadabur, a large-scale Quran audio dataset. Tadabur comprises more than 1400+ hours of…

Sound · Computer Science 2026-04-22 Faisal Alherran

In this paper, we introduce TEDxTN, the first publicly available Tunisian Arabic to English speech translation dataset. This work is in line with the ongoing effort to mitigate the data scarcity obstacle for a number of Arabic dialects. We…

Computation and Language · Computer Science 2025-11-17 Fethi Bougares , Salima Mdhaffar , Haroun Elleuch , Yannick Estève

The advancement of speech technologies has been remarkable, yet its integration with African languages remains limited due to the scarcity of African speech corpora. To address this issue, we present AfroDigits, a minimalist,…

Ethiopic/Amharic script is one of the oldest African writing systems, which serves at least 23 languages (e.g., Amharic, Tigrinya) in East Africa for more than 120 million people. The Amharic writing system, Abugida, has 282 syllables, 15…

Computer Vision and Pattern Recognition · Computer Science 2022-04-12 Wondimu Dikubab , Dingkang Liang , Minghui Liao , Xiang Bai

In this paper, we introduce SaudiBERT, a monodialect Arabic language model pretrained exclusively on Saudi dialectal text. To demonstrate the model's effectiveness, we compared SaudiBERT with six different multidialect Arabic language…

Computation and Language · Computer Science 2024-05-13 Faisal Qarah

Despite the importance of handwritten numeral classification, a robust and effective method for a widely used language like Arabic is still due. This study focuses to overcome two major limitations of existing works: data diversity and…

Computer Vision and Pattern Recognition · Computer Science 2019-08-07 S. M. A. Sharif , Ghulam Mujtaba , S. M. Nadim Uddin

Labelling of user's utterances to understanding his attends which called Dialogue Act (DA) classification, it is considered the key player for dialogue language understanding layer in automatic dialogue systems. In this paper, we proposed a…

Computation and Language · Computer Science 2015-09-11 Abdelrahim A Elmadany , Sherif M Abdou , Mervat Gheith

In this paper we present the final result of a project on Tunisian Arabic encoded in Arabizi, the Latin-based writing system for digital conversations. The project led to the creation of two integrated and independent resources: a corpus…

Computation and Language · Computer Science 2022-07-12 Elisa Gugliotta , Marco Dinarelli

This study presents systems submitted by the University of Texas at Dallas, Center for Robust Speech Systems (UTD-CRSS) to the MGB-3 Arabic Dialect Identification (ADI) subtask. This task is defined to discriminate between five dialects of…

Audio and Speech Processing · Electrical Eng. & Systems 2017-10-03 Ahmet E. Bulut , Qian Zhang , Chunlei Zhang , Fahimeh Bahmaninezhad , John H. L. Hansen

Arabic is a widely-spoken language with a long and rich history, but existing corpora and language technology focus mostly on modern Arabic and its varieties. Therefore, studying the history of the language has so far been mostly limited to…

Computation and Language · Computer Science 2018-09-12 Yonatan Belinkov , Alexander Magidow , Alberto Barrón-Cedeño , Avi Shmidman , Maxim Romanov

Commonsense validation evaluates whether a sentence aligns with everyday human understanding, a critical capability for developing robust natural language understanding systems. While substantial progress has been made in English, the task…

Computation and Language · Computer Science 2026-01-13 Kareem Elozeiri , Mervat Abassy , Preslav Nakov , Yuxia Wang

Lipreading involves using visual data to recognize spoken words by analyzing the movements of the lips and surrounding area. It is a hot research topic with many potential applications, such as human-machine interaction and enhancing audio…

Computer Vision and Pattern Recognition · Computer Science 2024-09-20 Samar Daou , Achraf Ben-Hamadou , Ahmed Rekik , Abdelaziz Kallel

We present the speech to text transcription system, called DARTS, for low resource Egyptian Arabic dialect. We analyze the following; transfer learning from high resource broadcast domain to low-resource dialectal domain and semi-supervised…

Computation and Language · Computer Science 2019-09-27 Sameer Khurana , Ahmed Ali , James Glass

Arabic spans over 30 spoken varieties, yet no open-source text-to-speech system unifies them. Key barriers include substantial cross-dialect lexical and phonological divergence, scarce synthesis-grade data, and the absence of a standardized…

Computation and Language · Computer Science 2026-04-01 Yushen Chen , Junzhe Liu , Yujie Tu , Zhikang Niu , Yuzhe Liang , Chunyu Qiang , Chen Zhang , Kai Yu , Xie Chen

Recently, the AI community has made significant strides in developing powerful foundation models, driven by large-scale multimodal datasets. However, for audio representation learning, existing datasets suffer from limitations in the…

Sound · Computer Science 2024-09-10 Luoyi Sun , Xuenan Xu , Mengyue Wu , Weidi Xie

Large language models (LLMs) for Arabic are still dominated by Modern Standard Arabic (MSA), with limited support for Saudi dialects such as Najdi and Hijazi. This underrepresentation hinders their ability to capture authentic dialectal…

Computation and Language · Computer Science 2025-08-20 Hassan Barmandah

In this paper, an approach for hate speech detection against women in Arabic community on social media (e.g. Youtube) is proposed. In the literature, similar works have been presented for other languages such as English. However, to the…

Computation and Language · Computer Science 2021-04-06 Imane Guellil , Ahsan Adeel , Faical Azouaou , Mohamed Boubred , Yousra Houichi , Akram Abdelhaq Moumna

This survey provides the first systematic review of Arabic LLM benchmarks, analyzing 40+ evaluation benchmarks across NLP tasks, knowledge domains, cultural understanding, and specialized capabilities. We propose a taxonomy organizing…

‹ Prev 1 4 5 6 7 8 10 Next ›