English
Related papers

Related papers: 101 Billion Arabic Words Dataset

200 papers

With over 2,000 languages and potentially millions of speakers, Africa represents one of the richest linguistic regions in the world. Yet, this diversity is scarcely reflected in state-of-the-art natural language processing (NLP) systems…

Computation and Language · Computer Science 2025-10-03 Jesujoba O. Alabi , Michael A. Hedderich , David Ifeoluwa Adelani , Dietrich Klakow

As the reach of large language models (LMs) expands globally, their ability to cater to diverse cultural contexts becomes crucial. Despite advancements in multilingual capabilities, models are not designed with appropriate cultural nuances.…

Computation and Language · Computer Science 2024-03-21 Tarek Naous , Michael J. Ryan , Alan Ritter , Wei Xu

Data sparsity is a main problem hindering the development of code-switching (CS) NLP systems. In this paper, we investigate data augmentation techniques for synthesizing dialectal Arabic-English CS text. We perform lexical replacements…

Computation and Language · Computer Science 2023-04-05 Injy Hamed , Nizar Habash , Slim Abdennadher , Ngoc Thang Vu

This paper presents the design and development of multi-dialect automatic speech recognition for Arabic. Deep neural networks are becoming an effective tool to solve sequential data problems, particularly, adopting an end-to-end training of…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-30 Abbas Raza Ali

Large Language Models (LLMs) have garnered remarkable advancements across diverse code-related tasks, known as Code LLMs, particularly in code generation that generates source code with LLM from natural language descriptions. This…

Computation and Language · Computer Science 2025-10-28 Juyong Jiang , Fan Wang , Jiasi Shen , Sungju Kim , Sunghun Kim

Instruction tuning has emerged as a prominent methodology for teaching Large Language Models (LLMs) to follow instructions. However, current instruction datasets predominantly cater to English or are derived from English-dominated LLMs,…

This paper adding more insights towards resources and datasets used in Arabic offensive language research. The main goal of this paper is to guide researchers in Arabic offensive language in selecting appropriate datasets based on their…

Computation and Language · Computer Science 2021-01-28 Fatemah Husain , Ozlem Uzuner

The last two years have seen a rapid growth in concerns around the safety of large language models (LLMs). Researchers and practitioners have met these concerns by creating an abundance of datasets for evaluating and improving LLM safety.…

Computation and Language · Computer Science 2025-01-13 Paul Röttger , Fabio Pernisi , Bertie Vidgen , Dirk Hovy

Motivated by the widespread increase in the phenomenon of code-switching between Egyptian Arabic and English in recent times, this paper explores the intricacies of machine translation (MT) and automatic speech recognition (ASR) systems,…

Computation and Language · Computer Science 2024-07-16 Ahmed Heakl , Youssef Zaghloul , Mennatullah Ali , Rania Hossam , Walid Gomaa

As large language models (LLMs) become increasingly central to Arabic NLP applications, evaluating their understanding of regional dialects and cultural nuances is essential, particularly in linguistically diverse settings like Saudi…

Computation and Language · Computer Science 2025-12-23 Renad Al-Monef , Hassan Alhuzali , Nora Alturayeif , Ashwag Alasmari

Despite representing nearly one-third of the world's languages, African languages remain critically underserved by modern NLP technologies, with 88\% classified as severely underrepresented or completely ignored in computational…

The largest dataset of Arabic speech mispronunciation detections in Egyptian dialogues is introduced. The dataset is composed of annotated audio files representing the top 100 words that are most frequently used in the Arabic language,…

Computation and Language · Computer Science 2021-11-03 Salah A. Aly , Abdelrahman Salah , Hesham M. Eraqi

This work investigates how effectively large language models (LLMs) and their tokenization schemes represent and generate Arabic root-pattern morphology, probing whether they capture genuine morphological structure or rely on surface…

Computation and Language · Computer Science 2026-03-18 Yara Alakeel , Chatrine Qwaider , Hanan Aldarmaki , Sawsan Alqahtani

In recent years, large language models (LLMs) have demonstrated significant potential across various natural language processing (NLP) tasks. However, their performance in domain-specific applications and non-English languages remains less…

Computation and Language · Computer Science 2025-10-01 Dragos-Dumitru Ghinea , Adela-Nicoleta Corbeanu , Adrian-Marius Dumitran

Sentiment analysis is a task of natural language processing which has recently attracted increasing attention. However, sentiment analysis research has mainly been carried out for the English language. Although Arabic is ramping up as one…

Computation and Language · Computer Science 2020-06-02 Oumaima Oueslati , Erik Cambria , Moez Ben HajHmida , Habib Ounelli

The rapid evolution of Natural Language Processing (NLP) has favoured major languages such as English, leaving a significant gap for many others due to limited resources. This is especially evident in the context of data annotation, a task…

Large language models (LLMs) perform strongly on many NLP tasks, but their ability to produce explicit linguistic structure remains unclear. We evaluate instruction-tuned LLMs on two structured prediction tasks for Standard Arabic:…

Computation and Language · Computer Science 2026-03-18 Mohamed Adel , Bashar Alhafni , Nizar Habash

Arabic text recognition is a challenging task because of the cursive nature of Arabic writing system, its joint writing scheme, the large number of ligatures and many other challenges. Deep Learning DL models achieved significant progress…

Computer Vision and Pattern Recognition · Computer Science 2020-09-07 Mohammad Fasha , Bassam Hammo , Nadim Obeid , Jabir Widian

In spite of the recent progress in speech processing, the majority of world languages and dialects remain uncovered. This situation only furthers an already wide technological divide, thereby hindering technological and socioeconomic…

Arabic language lacks semantic datasets and sense inventories. The most common semantically-labeled dataset for Arabic is the ArabGlossBERT, a relatively small dataset that consists of 167K context-gloss pairs (about 60K positive and 107K…

Computation and Language · Computer Science 2023-02-09 Sanad Malaysha , Mustafa Jarrar , Mohammed Khalilia