English
Related papers

Related papers: Camelira: An Arabic Multi-Dialect Morphological Di…

200 papers

Large Language Models (LLMs) have shown impressive results in multiple domains of natural language processing (NLP) but are mainly focused on the English language. Recently, more LLMs have incorporated a larger proportion of multilingual…

The rapid advancements in Large Language Models (LLMs) have led to significant improvements in various natural language processing tasks. However, the evaluation of LLMs' legal knowledge, particularly in non-English languages such as…

In this article we focus firstly on the principle of pedagogical indexing and characteristics of Arabic language and secondly on the possibility of adapting the standard for describing learning resources used (the LOM and its Application…

Computation and Language · Computer Science 2012-08-02 Asma Boudhief , Mohsen Maraoui , Mounir Zrigui

This paper addresses critical gaps in Arabic language model evaluation by establishing comprehensive theoretical guidelines and introducing a novel evaluation framework. We first analyze existing Arabic evaluation datasets, identifying…

Computation and Language · Computer Science 2025-06-03 Serry Sibaee , Omer Nacar , Adel Ammar , Yasser Al-Habashi , Abdulrahman Al-Batati , Wadii Boulila

Large language models (LLMs) have achieved remarkable progress in many language tasks, yet they continue to struggle with complex historical and religious Arabic texts such as the Quran and Hadith. To address this limitation, we develop a…

Computation and Language · Computer Science 2026-03-26 Somaya Eltanbouly , Samer Rashwani

Data contamination undermines the validity of Large Language Model evaluation by enabling models to rely on memorized benchmark content rather than true generalization. While prior work has proposed contamination detection methods, these…

Computation and Language · Computer Science 2026-01-22 Chaymaa Abbas , Nour Shamaa , Mariette Awad

Arabic remains one of the most underrepresented languages in natural language processing research, particularly in medical applications, due to the limited availability of open-source data and benchmarks. The lack of resources hinders…

Computation and Language · Computer Science 2026-02-03 Mouath Abu-Daoud , Leen Kharouf , Omar El Hajj , Dana El Samad , Mariam Al-Omari , Jihad Mallat , Khaled Saleh , Nizar Habash , Farah E. Shamout

The paper describes the Egyptian Arabic-to-English statistical machine translation (SMT) system that the QCRI-Columbia-NYUAD (QCN) group submitted to the NIST OpenMT'2015 competition. The competition focused on informal dialectal Arabic, as…

Computation and Language · Computer Science 2016-06-21 Hassan Sajjad , Nadir Durrani , Francisco Guzman , Preslav Nakov , Ahmed Abdelali , Stephan Vogel , Wael Salloum , Ahmed El Kholy , Nizar Habash

Automatic Arabic Dialect Identification (ADI) of text has gained great popularity since it was introduced in the early 2010s. Multiple datasets were developed, and yearly shared tasks have been running since 2018. However, ADI systems are…

Computation and Language · Computer Science 2023-10-23 Amr Keleg , Walid Magdy

Darija Open Dataset (DODa) is an open-source project for the Moroccan dialect. With more than 10,000 entries DODa is arguably the largest open-source collaborative project for Darija-English translation built for Natural Language Processing…

Computation and Language · Computer Science 2021-03-18 Aissam Outchakoucht , Hamza Es-Samaali

In derivational morphology, what mechanisms govern the variation in form-meaning relations between words? The answers to this type of questions are typically based on intuition and on observations drawn from limited data, even when a wide…

Computation and Language · Computer Science 2026-04-15 Hathout Nabil , Basilio Calderone , Fiammetta Namer , Franck Sajous

This paper describes the Arabic MGB-3 Challenge - Arabic Speech Recognition in the Wild. Unlike last year's Arabic MGB-2 Challenge, for which the recognition task was based on more than 1,200 hours broadcast TV news recordings from…

Computation and Language · Computer Science 2017-09-22 Ahmed Ali , Stephan Vogel , Steve Renals

Natural Language Processing (NLP) is a vital computational method for addressing language processing, analysis, and generation. NLP tasks form the core of many daily applications, from automatic text correction to speech recognition. While…

Computation and Language · Computer Science 2024-10-18 Caroline Sabty

SALMA, the first Arabic sense-annotated corpus, consists of ~34K tokens, which are all sense-annotated. The corpus is annotated using two different sense inventories simultaneously (Modern and Ghani). SALMA novelty lies in how tokens and…

Computation and Language · Computer Science 2023-10-31 Mustafa Jarrar , Sanad Malaysha , Tymaa Hammouda , Mohammed Khalilia

We introduce Atlas-Chat, the first-ever collection of LLMs specifically developed for dialectal Arabic. Focusing on Moroccan Arabic, also known as Darija, we construct our instruction dataset by consolidating existing Darija language…

Dialectal Arabic (DA) poses a persistent challenge for natural language processing (NLP), as most everyday communication in the Arab world occurs in dialects that diverge significantly from Modern Standard Arabic (MSA). This linguistic…

Computation and Language · Computer Science 2025-09-04 Abdullah Alabdullah , Lifeng Han , Chenghua Lin

Current authentication and trusted systems depend on classical and biometric methods to recognize or authorize users. Such methods include audio speech recognitions, eye, and finger signatures. Recent tools utilize deep learning and…

Sound · Computer Science 2021-11-12 Aly Moustafa , Salah A. Aly

Building dialogues systems interaction has recently gained considerable attention, but most of the resources and systems built so far are tailored to English and other Indo-European languages. The need for designing systems for other…

Computation and Language · Computer Science 2015-05-13 AbdelRahim A. Elmadany , Sherif M. Abdou , Mervat Gheith

In this paper, we present an ongoing effort in lexical semantic analysis and annotation of Modern Standard Arabic (MSA) text, a semi automatic annotation tool concerned with the morphologic, syntactic, and semantic levels of description.

Computation and Language · Computer Science 2016-05-25 Abdelaziz Lakhfif , Mohammed T. Laskri , Eric Atwell

This paper presents a comparative study for window-based descriptors on the application of Arabic handwritten alphabet recognition. We show a detailed experimental evaluation of different descriptors with several classifiers. The objective…

Computer Vision and Pattern Recognition · Computer Science 2014-11-18 Marwan Torki , Mohamed E. Hussein , Ahmed Elsallamy , Mahmoud Fayyaz , Shehab Yaser