English
Related papers

Related papers: NADI 2025: The First Multidialectal Arabic Speech …

200 papers

The Arabic language is among the most popular languages in the world with a huge variety of dialects spoken in 22 countries. In this study, we address the problem of classifying 18 Arabic dialects of the QADI dataset of Arabic tweets. RNN…

Computation and Language · Computer Science 2025-07-01 Omar A. Essameldin , Ali O. Elbeih , Wael H. Gomaa , Wael F. Elsersy

Arabic has diverse dialects, where one dialect can be substantially different from the others. In the NLP literature, some assumptions about these dialects are widely adopted (e.g., ``Arabic dialects can be grouped into distinguishable…

Computation and Language · Computer Science 2025-05-29 Amr Keleg , Sharon Goldwater , Walid Magdy

This paper describes the SemEval--2016 Task 3 on Community Question Answering, which we offered in English and Arabic. For English, we had three subtasks: Question--Comment Similarity (subtask A), Question--Question Similarity (B), and…

Computation and Language · Computer Science 2019-12-05 Preslav Nakov , Lluís Màrquez , Alessandro Moschitti , Walid Magdy , Hamdy Mubarak , Abed Alhakim Freihat , James Glass , Bilal Randeree

This study investigates logistic regression, linear support vector machine, multinomial Naive Bayes, and Bernoulli Naive Bayes for classifying Libyan dialect utterances gathered from Twitter. The dataset used is the QADI corpus, which…

Computation and Language · Computer Science 2025-12-05 Mansour Essgaer , Khamis Massud , Rabia Al Mamlook , Najah Ghmaid

The rapid spread of online disinformation presents a global challenge, and machine learning has been widely explored as a potential solution. However, multilingual settings and low-resource languages are often neglected in this field. To…

We describe SemEval-2017 Task 3 on Community Question Answering. This year, we reran the four subtasks from SemEval-2016:(A) Question-Comment Similarity,(B) Question-Question Similarity,(C) Question-External Comment Similarity, and (D)…

Computation and Language · Computer Science 2019-12-03 Preslav Nakov , Doris Hoogeveen , Lluís Màrquez , Alessandro Moschitti , Hamdy Mubarak , Timothy Baldwin , Karin Verspoor

Question semantic similarity (Q2Q) is a challenging task that is very useful in many NLP applications, such as detecting duplicate questions and question answering systems. In this paper, we present the results and findings of the shared…

Computation and Language · Computer Science 2019-09-24 Haitham Seelawi , Ahmad Mustafa , Hesham Al-Bataineh , Wael Farhan , Hussein T. Al-Natsheh

This paper addresses the problem of Bangla hate speech identification, a socially impactful yet linguistically challenging task. As part of the "Bangla Multi-task Hate Speech Identification" shared task at the BLP Workshop, IJCNLP-AACL…

Computation and Language · Computer Science 2025-11-11 Sourav Saha , K M Nafi Asib , Mohammed Moshiul Hoque

In this paper, we tackle the Arabic Fine-Grained Hate Speech Detection shared task and demonstrate significant improvements over reported baselines for its three subtasks. The tasks are to predict if a tweet contains (1) Offensive language;…

Computation and Language · Computer Science 2022-05-18 Badr AlKhamissi , Mona Diab

Arabic dialect identification is a specific task of natural language processing, aiming to automatically predict the Arabic dialect of a given text. Arabic dialect identification is the first step in various natural language processing…

Computation and Language · Computer Science 2020-09-29 Maha J. Althobaiti

Arabic dialect recognition presents a significant challenge in speech technology due to the linguistic diversity of Arabic and the scarcity of large annotated datasets, particularly for underrepresented dialects. This research investigates…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-27 Ghazal Al-Shwayyat , Omer Nezih Gerek

We present WojoodNER-2023, the first Arabic Named Entity Recognition (NER) Shared Task. The primary focus of WojoodNER-2023 is on Arabic NER, offering novel NER datasets (i.e., Wojood) and the definition of subtasks designed to facilitate…

Computation and Language · Computer Science 2023-10-26 Mustafa Jarrar , Muhammad Abdul-Mageed , Mohammed Khalilia , Bashar Talafha , AbdelRahim Elmadany , Nagham Hamad , Alaa' Omar

Dialectal Arabic (DA) speech data vary widely in domain coverage, dialect labeling practices, and recording conditions, complicating cross-dataset comparison and model evaluation. To characterize this landscape, we conduct a computational…

Computation and Language · Computer Science 2026-01-30 Peter Sullivan , AbdelRahim Elmadany , Alcides Alcoba Inciarte , Muhammad Abdul-Mageed

This work is an attempt to introduce a comprehensive benchmark for Arabic speech recognition, specifically tailored to address the challenges of telephone conversations in Arabic language. Arabic, characterized by its rich dialectal…

Artificial Intelligence · Computer Science 2024-05-31 Qusai Abo Obaidah , Muhy Eddin Za'ter , Adnan Jaljuli , Ali Mahboub , Asma Hakouz , Bashar Al-Rfooh , Yazan Estaitia

Prediction of language varieties and dialects is an important language processing task, with a wide range of applications. For Arabic, the native tongue of ~ 300 million people, most varieties remain unsupported. To ease this bottleneck, we…

Computation and Language · Computer Science 2019-11-01 Muhammad Abdul-Mageed , Chiyu Zhang , AbdelRahim Elmadany , Arun Rajendran , Lyle Ungar

This paper describes the Arabic MGB-3 Challenge - Arabic Speech Recognition in the Wild. Unlike last year's Arabic MGB-2 Challenge, for which the recognition task was based on more than 1,200 hours broadcast TV news recordings from…

Computation and Language · Computer Science 2017-09-22 Ahmed Ali , Stephan Vogel , Steve Renals

While language identification is a fundamental speech and language processing task, for many languages and language families it remains a challenging task. For many low-resource and endangered languages this is in part due to resource…

Discriminating between closely-related language varieties is considered a challenging and important task. This paper describes our submission to the DSL 2016 shared-task, which included two sub-tasks: one on discriminating similar languages…

Computation and Language · Computer Science 2016-09-27 Yonatan Belinkov , James Glass

Being modeled as a single-label classification task for a long time, recent work has argued that Arabic Dialect Identification (ADI) should be framed as a multi-label classification task. However, ADI remains constrained by the availability…

Computation and Language · Computer Science 2026-02-18 Ali Mekky , Mohamed El Zeftawy , Lara Hassan , Amr Keleg , Preslav Nakov