中文
相关论文

相关论文: Normalized Orthography for Tunisian Arabic

200 篇论文

Amharic is one of the official languages of the Federal Democratic Republic of Ethiopia. It is one of the languages that use an Ethiopic script which is derived from Gee'z, ancient and currently a liturgical language. Amharic is also one of…

计算机视觉与模式识别 · 计算机科学 2022-02-28 Mesay Samuel Gondere , Lars Schmidt-Thieme , Durga Prasad Sharma , Abiot Sinamo Boltena

The recognition of unconstrained handwriting continues to be a difficult task for computers despite active research for several decades. This is because handwritten text offers great challenges such as character and word segmentation,…

神经与进化计算 · 计算机科学 2013-01-22 Yusuf Perwej

Arabic Language Models (LMs) are pretrained predominately on Modern Standard Arabic (MSA) and are expected to transfer to its dialects. While MSA as the standard written variety is commonly used in formal settings, people speak and write…

计算与语言 · 计算机科学 2026-04-01 Abdulmuizz Khalak , Abderrahmane Issam , Gerasimos Spanakis

Automatic diacritization of Arabic text involves adding diacritical marks (diacritics) to the text. This task poses a significant challenge with noteworthy implications for computational processing and comprehension. In this paper, we…

计算与语言 · 计算机科学 2024-01-11 Abderrahman Skiredj , Ismail Berrada

The growing importance of culturally-aware natural language processing systems has led to an increasing demand for resources that capture sociopragmatic phenomena across diverse languages. Nevertheless, Arabic-language resources for…

In this paper we address the scarcity of annotated data for NArabizi, a Romanized form of North African Arabic used mostly on social media, which poses challenges for Natural Language Processing (NLP). We introduce an enriched version of…

计算与语言 · 计算机科学 2024-12-06 Arij Riabi , Menel Mahamdi , Djamé Seddah

A sufficient amount of annotated data is usually required to fine-tune pre-trained language models for downstream tasks. Unfortunately, attaining labeled data can be costly, especially for multiple language varieties and dialects. We…

计算与语言 · 计算机科学 2021-02-04 Muhammad Khalifa , Muhammad Abdul-Mageed , Khaled Shaalan

Arabic is a semitic language characterized by a complex and rich morphology. The exceptional degree of ambiguity in the writing system, the rich morphology, and the highly complex word formation process of roots and patterns all contribute…

计算机视觉与模式识别 · 计算机科学 2014-12-25 Ibrahim Abdelaziz , Sherif Abdou

This paper presents an Arabic Alphabet Sign Language recognition approach, using deep learning methods in conjunction with transfer learning and transformer-based models. We study the performance of the different variants on two publicly…

计算机视觉与模式识别 · 计算机科学 2024-10-02 Mazen Balat , Rewaa Awaad , Hend Adel , Ahmed B. Zaky , Salah A. Aly

Arabic spans over 30 spoken varieties, yet no open-source text-to-speech system unifies them. Key barriers include substantial cross-dialect lexical and phonological divergence, scarce synthesis-grade data, and the absence of a standardized…

计算与语言 · 计算机科学 2026-04-01 Yushen Chen , Junzhe Liu , Yujie Tu , Zhikang Niu , Yuzhe Liang , Chunyu Qiang , Chen Zhang , Kai Yu , Xie Chen

Dialect and standard language identification are crucial tasks for many Arabic natural language processing applications. In this paper, we present our deep learning-based system, submitted to the second NADI shared task for country-level…

计算与语言 · 计算机科学 2021-06-24 Abdellah El Mekki , Abdelkader El Mahdaouy , Kabil Essefar , Nabil El Mamoun , Ismail Berrada , Ahmed Khoumsi

Neural machine translation (NMT) has shown impressive performance when trained on large-scale corpora. However, generic NMT systems have demonstrated poor performance on out-of-domain translation. To mitigate this issue, several domain…

计算与语言 · 计算机科学 2023-09-25 Emad A. Alghamdi , Jezia Zakraoui , Fares A. Abanmy

We study the performance of Arabic text classification combining various techniques: (a) tfidf vs. dependency syntax, for feature selection and weighting; (b) class association rules vs. support vector machines, for classification. The…

计算与语言 · 计算机科学 2014-10-21 Yannis Haralambous , Yassir Elidrissi , Philippe Lenca

Transcribed speech and user-generated text in Arabic typically contain a mixture of Modern Standard Arabic (MSA), the standardized language taught in schools, and Dialectal Arabic (DA), used in daily communications. To handle this…

计算与语言 · 计算机科学 2023-10-24 Amr Keleg , Sharon Goldwater , Walid Magdy

The rich linguistic landscape of the Arab world is characterized by a significant gap between Modern Standard Arabic (MSA), the language of formal communication, and the diverse regional dialects used in everyday life. This diglossia…

计算与语言 · 计算机科学 2025-08-05 Serry Sibaee , Omer Nacar , Yasser Al-Habashi , Adel Ammar , Wadii Boulila

Designing a natural voice interface rely mostly on Speech recognition for interaction between human and their modern digital life equipment. In addition, speech recognition narrows the gap between monolingual individuals to better exchange…

计算与语言 · 计算机科学 2022-12-22 Ayman Mansour , Wafaa F. Mukhtar

SALMA, the first Arabic sense-annotated corpus, consists of ~34K tokens, which are all sense-annotated. The corpus is annotated using two different sense inventories simultaneously (Modern and Ghani). SALMA novelty lies in how tokens and…

计算与语言 · 计算机科学 2023-10-31 Mustafa Jarrar , Sanad Malaysha , Tymaa Hammouda , Mohammed Khalilia

Darija Open Dataset (DODa) is an open-source project for the Moroccan dialect. With more than 10,000 entries DODa is arguably the largest open-source collaborative project for Darija-English translation built for Natural Language Processing…

计算与语言 · 计算机科学 2021-03-18 Aissam Outchakoucht , Hamza Es-Samaali

The rapid evolution of Natural Language Processing (NLP) has favoured major languages such as English, leaving a significant gap for many others due to limited resources. This is especially evident in the context of data annotation, a task…

We present Camelira, a web-based Arabic multi-dialect morphological disambiguation tool that covers four major variants of Arabic: Modern Standard Arabic, Egyptian, Gulf, and Levantine. Camelira offers a user-friendly web interface that…

计算与语言 · 计算机科学 2022-12-01 Ossama Obeid , Go Inoue , Nizar Habash