English
Related papers

Related papers: Lisan: Yemeni, Iraqi, Libyan, and Sudanese Arabic …

200 papers

In this paper, we present Arap-Tweet, which is a large-scale and multi-dialectal corpus of Tweets from 11 regions and 16 countries in the Arab world representing the major Arabic dialectal varieties. To build this corpus, we collected data…

Computation and Language · Computer Science 2018-08-24 Wajdi Zaghouani , Anis Charfi

We present the QuranMorph corpus, a morphologically annotated corpus for the Quran (77,429 tokens). Each token in the QuranMorph was manually lemmatized and tagged with its part-of-speech by three expert linguists. The lemmatization process…

Computation and Language · Computer Science 2025-06-24 Diyam Akra , Tymaa Hammouda , Mustafa Jarrar

This paper presents Nabra, a corpora of Syrian Arabic dialects with morphological annotations. A team of Syrian natives collected more than 6K sentences containing about 60K words from several sources including social media posts, scripts…

Computation and Language · Computer Science 2023-10-27 Amal Nayouf , Tymaa Hammouda , Mustafa Jarrar , Fadi Zaraket , Mohamad-Bassam Kurdy

In this paper, we present the annotation pipeline and the guidelines we wrote as part of an effort to create a large manually annotated Arabic author profiling dataset from various social media sources covering 16 Arabic countries and 11…

Computation and Language · Computer Science 2018-08-24 Wajdi Zaghouani , Anis Charfi

We present our effort to create a large Multi-Layered representational repository of Linguistic Code-Switched Arabic data. The process involves developing clear annotation standards and Guidelines, streamlining the annotation process, and…

Computation and Language · Computer Science 2019-10-01 Mona Diab , Mahmoud Ghoneim , Abdelati Hawwari , Fahad AlGhamdi , Nada AlMarwani , Mohamed Al-Badrashiny

SALMA, the first Arabic sense-annotated corpus, consists of ~34K tokens, which are all sense-annotated. The corpus is annotated using two different sense inventories simultaneously (Modern and Ghani). SALMA novelty lies in how tokens and…

Computation and Language · Computer Science 2023-10-31 Mustafa Jarrar , Sanad Malaysha , Tymaa Hammouda , Mohammed Khalilia

We present QADI, an automatically collected dataset of tweets belonging to a wide range of country-level Arabic dialects -covering 18 different countries in the Middle East and North Africa region. Our method for building this dataset…

Computation and Language · Computer Science 2020-05-18 Ahmed Abdelali , Hamdy Mubarak , Younes Samih , Sabit Hassan , Kareem Darwish

This study investigates logistic regression, linear support vector machine, multinomial Naive Bayes, and Bernoulli Naive Bayes for classifying Libyan dialect utterances gathered from Twitter. The dataset used is the QADI corpus, which…

Computation and Language · Computer Science 2025-12-05 Mansour Essgaer , Khamis Massud , Rabia Al Mamlook , Najah Ghmaid

Text summarization has been intensively studied in many languages, and some languages have reached advanced stages. Yet, Arabic Text Summarization (ATS) is still in its developing stages. Existing ATS datasets are either small or lack…

Computation and Language · Computer Science 2022-10-26 Abdulaziz Alhamadani , Xuchao Zhang , Jianfeng He , Chang-Tien Lu

Large language models (LLMs) for Arabic are still dominated by Modern Standard Arabic (MSA), with limited support for Saudi dialects such as Najdi and Hijazi. This underrepresentation hinders their ability to capture authentic dialectal…

Computation and Language · Computer Science 2025-08-20 Hassan Barmandah

The Arabic language is characterized by a rich tapestry of regional dialects that differ substantially in phonetics and lexicon, reflecting the geographic and cultural diversity of its speakers. Despite the availability of many…

The growing importance of culturally-aware natural language processing systems has led to an increasing demand for resources that capture sociopragmatic phenomena across diverse languages. Nevertheless, Arabic-language resources for…

The processing of the Arabic language is a complex field of research. This is due to many factors, including the complex and rich morphology of Arabic, its high degree of ambiguity, and the presence of several regional varieties that need…

Computation and Language · Computer Science 2022-05-20 Karim El Haff , Mustafa Jarrar , Tymaa Hammouda , Fadi Zaraket

We present Qabas, a novel open-source Arabic lexicon designed for NLP applications. The novelty of Qabas lies in its synthesis of 110 lexicons. Specifically, Qabas lexical entries (lemmas) are assembled by linking lemmas from 110 lexicons.…

Computation and Language · Computer Science 2024-06-12 Mustafa Jarrar , Tymaa Hammouda

The Arabic language is among the most popular languages in the world with a huge variety of dialects spoken in 22 countries. In this study, we address the problem of classifying 18 Arabic dialects of the QADI dataset of Arabic tweets. RNN…

Computation and Language · Computer Science 2025-07-01 Omar A. Essameldin , Ali O. Elbeih , Wael H. Gomaa , Wael F. Elsersy

Labelling of user's utterances to understanding his attends which called Dialogue Act (DA) classification, it is considered the key player for dialogue language understanding layer in automatic dialogue systems. In this paper, we proposed a…

Computation and Language · Computer Science 2015-09-11 Abdelrahim A Elmadany , Sherif M Abdou , Mervat Gheith

Identifying hate speech content in the Arabic language is challenging due to the rich quality of dialectal variations. This study introduces a multilabel hate speech dataset in the Arabic language. We have collected 10000 Arabic tweets and…

Computation and Language · Computer Science 2025-05-26 Wajdi Zaghouani , Md. Rafiul Biswas

Arabic is one of the most important and growing languages in the world. With the rise of social media platforms such as Twitter, Arabic spoken dialects have become more in use. In this paper, we describe our approach on the NADI Shared Task…

Computation and Language · Computer Science 2020-11-16 Ahmad Beltagy , Abdelrahman Wael , Omar ElSherief

In the context of low-resource languages, the Algerian dialect (AD) faces challenges due to the absence of annotated corpora, hindering its effective processing, notably in Machine Learning (ML) applications reliant on corpora for training…

Computation and Language · Computer Science 2024-11-08 Amin Abdedaiem , Abdelhalim Hafedh Dahou , Mohamed Amine Cheragui , Brigitte Mathiak

Most Arabic natural language processing tools and resources are developed to serve Modern Standard Arabic (MSA), which is the official written language in the Arab World. Some Dialectal Arabic varieties, notably Egyptian Arabic, have…

Computation and Language · Computer Science 2016-09-13 Salam Khalifa , Nizar Habash , Dana Abdulrahim , Sara Hassan
‹ Prev 1 2 3 10 Next ›