English
Related papers

Related papers: MIT-QCRI Arabic Dialect Identification System for …

200 papers

Deaf people are using sign language for communication, and it is a combination of gestures, movements, postures, and facial expressions that correspond to alphabets and words in spoken languages. The proposed Arabic sign language…

Computer Vision and Pattern Recognition · Computer Science 2023-01-10 Muhammad Al-Barham , Ahmad Jamal , Musa Al-Yaman

Despite the success of deep learning in speech recognition, multi-dialect speech recognition remains a difficult problem. Although dialect-specific acoustic models are known to perform well in general, they are not easy to maintain when…

Machine Learning · Computer Science 2022-05-09 Sanghyun Yoo , Inchul Song , Yoshua Bengio

Recent advancements have significantly enhanced the capabilities of Multimodal Large Language Models (MLLMs) in generating and understanding image-to-text content. Despite these successes, progress is predominantly limited to English due to…

Computation and Language · Computer Science 2024-07-29 Fakhraddin Alwajih , Gagan Bhatia , Muhammad Abdul-Mageed

This paper presents our approach to address the EACL WANLP-2021 Shared Task 1: Nuanced Arabic Dialect Identification (NADI). The task is aimed at developing a system that identifies the geographical location(country/province) from where an…

Computation and Language · Computer Science 2021-02-23 Anshul Wadhawan

Large Language Models (LLMs) have demonstrated significant promise for various applications in healthcare. However, their efficacy in the Arabic medical domain remains unexplored due to the lack of high-quality domain-specific datasets and…

Computation and Language · Computer Science 2025-08-25 Mouath Abu Daoud , Chaimae Abouzahir , Leen Kharouf , Walid Al-Eisawi , Nizar Habash , Farah E. Shamout

Recent advances in automatic speech recognition (ASR) have achieved accuracy levels comparable to human transcribers, which led researchers to debate if the machine has reached human performance. Previous work focused on the English…

Audio and Speech Processing · Electrical Eng. & Systems 2021-06-30 Amir Hussein , Shinji Watanabe , Ahmed Ali

Despite Arabic being one of the most widely spoken languages, the development of Arabic Automatic Speech Recognition (ASR) systems faces significant challenges due to the language's complexity, and only a limited number of public Arabic ASR…

Computation and Language · Computer Science 2025-07-21 Lilit Grigoryan , Nikolay Karpov , Enas Albasiri , Vitaly Lavrukhin , Boris Ginsburg

Deepfake generation methods are evolving fast, making fake media harder to detect and raising serious societal concerns. Most deepfake detection and dataset creation research focuses on monolingual content, often overlooking the challenges…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Kartik Kuckreja , Parul Gupta , Injy Hamed , Thamar Solorio , Muhammad Haris Khan , Abhinav Dhall

In this paper, we tackle the Arabic Fine-Grained Hate Speech Detection shared task and demonstrate significant improvements over reported baselines for its three subtasks. The tasks are to predict if a tweet contains (1) Offensive language;…

Computation and Language · Computer Science 2022-05-18 Badr AlKhamissi , Mona Diab

Recent advancements in zero-shot speech generation have enabled models to synthesize speech that mimics speaker identity and speaking style from speech prompts. However, these models' effectiveness is significantly limited in real-world…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-14 Boyu Zhu , Cheng Gong , Muyang Wu , Ruihao Jing , Fan Liu , Xiaolei Zhang , Chi Zhang , Xuelong Li

This paper introduces a new application named ArPA for Arabic kids who have trouble with pronunciation. Our application comprises two key components: the diagnostic module and the therapeutic module. The diagnostic process involves…

Arabic word segmentation is essential for a variety of NLP applications such as machine translation and information retrieval. Segmentation entails breaking words into their constituent stems, affixes and clitics. In this paper, we compare…

Computation and Language · Computer Science 2017-08-22 Mohamed Eldesouki , Younes Samih , Ahmed Abdelali , Mohammed Attia , Hamdy Mubarak , Kareem Darwish , Kallmeyer Laura

Diacritization process attempt to restore the short vowels in Arabic written text; which typically are omitted. This process is essential for applications such as Text-to-Speech (TTS). While diacritization of Modern Standard Arabic (MSA)…

Computation and Language · Computer Science 2019-06-03 Ahmed Abdelali , Mohammed Attia , Younes Samih , Kareem Darwish , Hamdy Mubarak

The prominence of figurative language devices, such as sarcasm and irony, poses serious challenges for Arabic Sentiment Analysis (SA). While previous research works tackle SA and sarcasm detection separately, this paper introduces an…

Computation and Language · Computer Science 2021-06-24 Abdelkader El Mahdaouy , Abdellah El Mekki , Kabil Essefar , Nabil El Mamoun , Ismail Berrada , Ahmed Khoumsi

Handwritten numeral recognition is in general a benchmark problem of Pattern Recognition and Artificial Intelligence. Compared to the problem of printed numeral recognition, the problem of handwritten numeral recognition is compounded due…

Computer Vision and Pattern Recognition · Computer Science 2010-03-10 Nibaran Das , Ayatullah Faruk Mollah , Sudip Saha , Syed Sahidul Haque

Neural machine translation (NMT) has shown impressive performance when trained on large-scale corpora. However, generic NMT systems have demonstrated poor performance on out-of-domain translation. To mitigate this issue, several domain…

Computation and Language · Computer Science 2023-09-25 Emad A. Alghamdi , Jezia Zakraoui , Fares A. Abanmy

Word embeddings are a core component of modern natural language processing systems, making the ability to thoroughly evaluate them a vital task. We describe DiaLex, a benchmark for intrinsic evaluation of dialectal Arabic word embedding.…

Arabic Text-to-Speech (TTS) research has been hindered by the availability of both publicly available training data and accurate Arabic diacritization models. In this paper, we address the limitation by exploring Arabic TTS training on…

Computation and Language · Computer Science 2026-03-03 Ahmed Musleh , Yifan Zhang , Kareem Darwish

Grammatical error correction (GEC) is a well-explored problem in English with many existing models and datasets. However, research on GEC in morphologically rich languages has been limited due to challenges such as data scarcity and…

Computation and Language · Computer Science 2023-11-10 Bashar Alhafni , Go Inoue , Christian Khairallah , Nizar Habash

Pre-trained Language Models (PLMs) are integral to many modern natural language processing (NLP) systems. Although multilingual models cover a wide range of languages, they often grapple with challenges like high inference costs and a lack…

Computation and Language · Computer Science 2024-07-19 Murtadha Ahmed , Saghir Alfasly , Bo Wen , Jamaal Qasem , Mohammed Ahmed , Yunfeng Liu