English
Related papers

Related papers: Creating a Large Multi-Layered Representational Re…

200 papers

In this paper, we present Arap-Tweet, which is a large-scale and multi-dialectal corpus of Tweets from 11 regions and 16 countries in the Arab world representing the major Arabic dialectal varieties. To build this corpus, we collected data…

Computation and Language · Computer Science 2018-08-24 Wajdi Zaghouani , Anis Charfi

In this paper, we present the annotation pipeline and the guidelines we wrote as part of an effort to create a large manually annotated Arabic author profiling dataset from various social media sources covering 16 Arabic countries and 11…

Computation and Language · Computer Science 2018-08-24 Wajdi Zaghouani , Anis Charfi

Natural Language Processing (NLP) is today a very active field of research and innovation. Many applications need however big sets of data for supervised learning, suitably labelled for the training purpose. This includes applications for…

Computation and Language · Computer Science 2021-02-23 ElMehdi Boujou , Hamza Chataoui , Abdellah El Mekki , Saad Benjelloun , Ikram Chairi , Ismail Berrada

Identifying hate speech content in the Arabic language is challenging due to the rich quality of dialectal variations. This study introduces a multilabel hate speech dataset in the Arabic language. We have collected 10000 Arabic tweets and…

Computation and Language · Computer Science 2025-05-26 Wajdi Zaghouani , Md. Rafiul Biswas

The growing importance of culturally-aware natural language processing systems has led to an increasing demand for resources that capture sociopragmatic phenomena across diverse languages. Nevertheless, Arabic-language resources for…

Code-switching is the phenomenon by which bilingual speakers switch between multiple languages during communication. The importance of developing language technologies for codeswitching data is immense, given the large populations that…

Computation and Language · Computer Science 2017-03-27 Victor Soto , Julia Hirschberg

When building NLP models, there is a tendency to aim for broader coverage, often overlooking cultural and (socio)linguistic nuance. In this position paper, we make the case for care and attention to such nuances, particularly in dataset…

Computation and Language · Computer Science 2022-03-21 A. Stevie Bergman , Mona T. Diab

The NLP pipeline has evolved dramatically in the last few years. The first step in the pipeline is to find suitable annotated datasets to evaluate the tasks we are trying to solve. Unfortunately, most of the published datasets lack metadata…

Computation and Language · Computer Science 2021-10-14 Zaid Alyafeai , Maraim Masoud , Mustafa Ghaleb , Maged S. Al-shaibani

The Arabic language is among the most popular languages in the world with a huge variety of dialects spoken in 22 countries. In this study, we address the problem of classifying 18 Arabic dialects of the QADI dataset of Arabic tweets. RNN…

Computation and Language · Computer Science 2025-07-01 Omar A. Essameldin , Ali O. Elbeih , Wael H. Gomaa , Wael F. Elsersy

We introduce MyVoice, a crowdsourcing platform designed to collect Arabic speech to enhance dialectal speech technologies. This platform offers an opportunity to design large dialectal speech datasets; and makes them publicly available.…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-08 Yousseif Elshahawy , Yassine El Kheir , Shammur Absar Chowdhury , Ahmed Ali

Natural Language Processing (NLP) is a vital computational method for addressing language processing, analysis, and generation. NLP tasks form the core of many daily applications, from automatic text correction to speech recognition. While…

Computation and Language · Computer Science 2024-10-18 Caroline Sabty

This survey provides the first systematic review of Arabic LLM benchmarks, analyzing 40+ evaluation benchmarks across NLP tasks, knowledge domains, cultural understanding, and specialized capabilities. We propose a taxonomy organizing…

We present an annotation schema as part of an effort to create a manually annotated corpus for Arabic dialogue language understanding including spoken dialogue and written "chat" dialogue for inquiry-answer domain. The proposed schema…

Computation and Language · Computer Science 2015-05-19 AbdelRahim A. Elmadany , Sherif M. Abdou , Mervat Gheith

Developing robust automatic speech recognition (ASR) systems for Arabic requires effective strategies to manage its diversity. Existing ASR systems mainly cover the modern standard Arabic (MSA) variety and few high-resource dialects, but…

Computation and Language · Computer Science 2025-06-02 Amirbek Djanibekov , Hawau Olamide Toyin , Raghad Alshalan , Abdullah Alitr , Hanan Aldarmaki

This article presents morphologically-annotated Yemeni, Sudanese, Iraqi, and Libyan Arabic dialects Lisan corpora. Lisan features around 1.2 million tokens. We collected the content of the corpora from several social media platforms. The…

Computation and Language · Computer Science 2022-12-20 Mustafa Jarrar , Fadi A Zaraket , Tymaa Hammouda , Daanish Masood Alavi , Martin Waahlisch

Understanding Arabic text and generating human-like responses is a challenging endeavor. While many researchers have proposed models and solutions for individual problems, there is an acute shortage of a comprehensive Arabic natural…

Computation and Language · Computer Science 2023-10-26 AbdelRahim Elmadany , El Moatez Billah Nagoudi , Muhammad Abdul-Mageed

Language in the Arab world presents a complex diglossic and multilingual setting, involving the use of Modern Standard Arabic, various dialects and sub-dialects, as well as multiple European languages. This diverse linguistic landscape has…

Computation and Language · Computer Science 2025-01-24 Injy Hamed , Caroline Sabty , Slim Abdennadher , Ngoc Thang Vu , Thamar Solorio , Nizar Habash

Transfer learning with a unified Transformer framework (T5) that converts all language problems into a text-to-text format was recently proposed as a simple and effective transfer learning approach. Although a multilingual version of the T5…

Computation and Language · Computer Science 2022-03-16 El Moatez Billah Nagoudi , AbdelRahim Elmadany , Muhammad Abdul-Mageed

The importance of building sentiment analysis tools for Arabic social media has been recognized during the past couple of years, especially with the rapid increase in the number of Arabic social media users. One of the main difficulties in…

Computation and Language · Computer Science 2017-10-26 Samhaa R. El-Beltagy , Talaat Khalil , Amal Halaby , Muhammad Hammad

We present QADI, an automatically collected dataset of tweets belonging to a wide range of country-level Arabic dialects -covering 18 different countries in the Middle East and North Africa region. Our method for building this dataset…

Computation and Language · Computer Science 2020-05-18 Ahmed Abdelali , Hamdy Mubarak , Younes Samih , Sabit Hassan , Kareem Darwish
‹ Prev 1 2 3 10 Next ›