中文
相关论文

相关论文: Arabic Dialect Identification in the Wild

200 篇论文

The largest dataset of Arabic speech mispronunciation detections in Egyptian dialogues is introduced. The dataset is composed of annotated audio files representing the top 100 words that are most frequently used in the Arabic language,…

计算与语言 · 计算机科学 2021-11-03 Salah A. Aly , Abdelrahman Salah , Hesham M. Eraqi

Arabic dialects have long been under-represented in Natural Language Processing (NLP) research due to their non-standardization and high variability, which pose challenges for computational modeling. Recent advances in the field, such as…

计算与语言 · 计算机科学 2026-02-19 Jonathan Mutal , Perla Al Almaoui , Simon Hengchen , Pierrette Bouillon

There is a significant gap in evaluating cultural reasoning in LLMs using conversational datasets that capture culturally rich and dialectal contexts. Most Arabic benchmarks focus on short text snippets in Modern Standard Arabic (MSA),…

Arabic is a linguistically and culturally rich language with a vast vocabulary that spans scientific, religious, and literary domains. Yet, large-scale lexical datasets linking Arabic words to precise definitions remain limited. We present…

计算与语言 · 计算机科学 2026-01-30 Serry Sibaee , Yasser Alhabashi , Nadia Sibai , Yara Farouk , Adel Ammar , Sawsan AlHalawani , Wadii Boulila

Language comprehension and commonsense knowledge validation by machines are challenging tasks that are still under researched and evaluated for Arabic text. In this paper, we present a benchmark Arabic dataset for commonsense explanation.…

计算与语言 · 计算机科学 2020-12-21 Saja AL-Tawalbeh , Mohammad AL-Smadi

We introduce the Tarab Corpus, a large-scale cultural and linguistic resource that brings together Arabic song lyrics and poetry within a unified analytical framework. The corpus comprises 2.56 million verses and more than 13.5 million…

计算与语言 · 计算机科学 2026-03-18 Mo El-Haj

This paper addresses the problem of detecting the offensive and abusive content in Facebook comments, where we focus on the Algerian dialectal Arabic which is one of under-resourced languages. The latter has a variety of dialects mixed with…

计算与语言 · 计算机科学 2022-03-21 Oussama Boucherit , Kheireddine Abainia

This paper proposes a novel approach to an automatic estimation of three speaker traits from Arabic speech: gender, emotion, and dialect. After showing promising results on different text classification tasks, the multi-task learning (MTL)…

计算与语言 · 计算机科学 2020-12-15 Wael Farhan , Muhy Eddin Za'ter , Qusai Abu Obaidah , Hisham al Bataineh , Zyad Sober , Hussein T. Al-Natsheh

Word embeddings are a core component of modern natural language processing systems, making the ability to thoroughly evaluate them a vital task. We describe DiaLex, a benchmark for intrinsic evaluation of dialectal Arabic word embedding.…

Question answering (QA) systems are now available through numerous commercial applications for a wide variety of domains, serving millions of users that interact with them via speech interfaces. However, current benchmarks in QA research do…

计算与语言 · 计算机科学 2021-09-27 Fahim Faisal , Sharlina Keshava , Md Mahfuz ibn Alam , Antonios Anastasopoulos

Arabic dialect identification (ADI) tools are an important part of the large-scale data collection pipelines necessary for training speech recognition models. As these pipelines require application of ADI tools to potentially out-of-domain…

音频与语音处理 · 电气工程与系统科学 2023-06-07 Peter Sullivan , AbdelRahim Elmadany , Muhammad Abdul-Mageed

We describe the findings of the fifth Nuanced Arabic Dialect Identification Shared Task (NADI 2024). NADI's objective is to help advance SoTA Arabic NLP by providing guidance, datasets, modeling opportunities, and standardized evaluation…

We describe AraNet, a collection of deep learning Arabic social media processing tools. Namely, we exploit an extensive host of publicly available and novel social media datasets to train bidirectional encoders from transformer models…

计算与语言 · 计算机科学 2020-04-14 Muhammad Abdul-Mageed , Chiyu Zhang , Azadeh Hashemi , El Moatez Billah Nagoudi

Assessing spoken language is challenging, and quantifying pronunciation metrics for machine learning models is even harder. However, for the Holy Quran, this task is simplified by the rigorous recitation rules (tajweed) established by…

音频与语音处理 · 电气工程与系统科学 2025-09-03 Abdullah Abdelfattah , Mahmoud I. Khalil , Hazem Abbas

Designing a natural voice interface rely mostly on Speech recognition for interaction between human and their modern digital life equipment. In addition, speech recognition narrows the gap between monolingual individuals to better exchange…

计算与语言 · 计算机科学 2022-12-22 Ayman Mansour , Wafaa F. Mukhtar

Automated Essay Scoring (AES) has gained increasing attention in recent years, yet research on Arabic AES remains limited due to the lack of publicly available datasets. To address this, we introduce LAILA, the largest publicly available…

With social media datasets being increasingly shared by researchers, it also presents the caveat that those datasets are not always completely replicable. Having to adhere to requirements of platforms like Twitter, researchers cannot…

数字图书馆 · 计算机科学 2018-03-08 Arkaitz Zubiaga

Although the prediction of dialects is an important language processing task, with a wide range of applications, existing work is largely limited to coarse-grained varieties. Inspired by geolocation research, we propose the novel task of…

计算与语言 · 计算机科学 2020-12-08 Muhammad Abdul-Mageed , Chiyu Zhang , AbdelRahim Elmadany , Lyle Ungar

With the rise of generative text-to-speech models, distinguishing between real and synthetic speech has become challenging, especially for Arabic that have received limited research attention. Most spoof detection efforts have focused on…

计算与语言 · 计算机科学 2025-09-30 Mohamed Maged , Alhassan Ehab , Ali Mekky , Besher Hassan , Shady Shehata

We have developed a system that automatically detects online jihadist hate speech with over 80% accuracy, by using techniques from Natural Language Processing and Machine Learning. The system is trained on a corpus of 45,000 subversive…

计算与语言 · 计算机科学 2018-03-14 Tom De Smedt , Guy De Pauw , Pieter Van Ostaeyen