English
Related papers

Related papers: New Arabic Medical Dataset for Diseases Classifica…

200 papers

Arabic text recognition is a challenging task because of the cursive nature of Arabic writing system, its joint writing scheme, the large number of ligatures and many other challenges. Deep Learning DL models achieved significant progress…

Computer Vision and Pattern Recognition · Computer Science 2020-09-07 Mohammad Fasha , Bassam Hammo , Nadim Obeid , Jabir Widian

Multilingual OCR and information extraction from receipts remains challenging, particularly for complex scripts like Arabic. We introduce \dataset, a comprehensive dataset designed for Arabic-English receipt understanding comprising 20,000…

We present ADI-20, an extension of the previously published ADI-17 Arabic Dialect Identification (ADI) dataset. ADI-20 covers all Arabic-speaking countries' dialects. It comprises 3,556 hours from 19 Arabic dialects in addition to Modern…

Computation and Language · Computer Science 2025-11-14 Haroun Elleuch , Salima Mdhaffar , Yannick Estève , Fethi Bougares

In spite of the recent progress in speech processing, the majority of world languages and dialects remain uncovered. This situation only furthers an already wide technological divide, thereby hindering technological and socioeconomic…

High-quality WordNets are crucial for achieving high-quality results in NLP applications that rely on such resources. However, the wordnets of most languages suffer from serious issues of correctness and completeness with respect to the…

Computation and Language · Computer Science 2024-04-01 Abed Alhakim Freihat , Hadi Khalilia , Gábor Bella , Fausto Giunchiglia

The problem of online offensive language limits the health and security of online users. It is essential to apply the latest state-of-the-art techniques in developing a system to detect online offensive language and to ensure social justice…

Computation and Language · Computer Science 2022-03-08 Fatemah Husain , Ozlem Uzuner

In this paper, we make freely accessible ANETAC our English-Arabic named entity transliteration and classification dataset that we built from freely available parallel translation corpora. The dataset contains 79,924 instances, each…

Computation and Language · Computer Science 2019-07-09 Mohamed Seghir Hadj Ameur , Farid Meziane , Ahmed Guessoum

This survey offers a comprehensive overview of Large Language Models (LLMs) designed for Arabic language and its dialects. It covers key architectures, including encoder-only, decoder-only, and encoder-decoder models, along with the…

Computation and Language · Computer Science 2026-05-20 Malak Mashaabi , Shahad Al-Khalifa , Hend Al-Khalifa

Classical and some deep learning techniques for Arabic text classification often depend on complex morphological analysis, word segmentation, and hand-crafted feature engineering. These could be eliminated by using character-level features.…

Computation and Language · Computer Science 2020-06-23 Mahmoud Daif , Shunsuke Kitada , Hitoshi Iyatomi

Text classification is an important task in Natural Language Processing (NLP), where the goal is to categorize text data into predefined classes. In this study, we analyse the dataset creation steps and evaluation techniques of multi-label…

Computation and Language · Computer Science 2023-03-01 Elmurod Kuriyozov , Ulugbek Salaev , Sanatbek Matlatipov , Gayrat Matlatipov

In this paper, we address the significant gap in Arabic natural language processing (NLP) resources by introducing ArabicaQA, the first large-scale dataset for machine reading comprehension and open-domain question answering in Arabic. This…

Computation and Language · Computer Science 2024-03-27 Abdelrahman Abdallah , Mahmoud Kasem , Mahmoud Abdalla , Mohamed Mahmoud , Mohamed Elkasaby , Yasser Elbendary , Adam Jatowt

Inaccuracies in existing or generated clinical text may lead to serious adverse consequences, especially if it is a misdiagnosis or incorrect treatment suggestion. With Large Language Models (LLMs) increasingly being used across diverse…

Computation and Language · Computer Science 2026-02-06 Congbo Ma , Yichun Zhang , Yousef Al-Jazzazi , Ahamed Foisal , Laasya Sharma , Yousra Sadqi , Khaled Saleh , Jihad Mallat , Farah E. Shamout

This paper presents the design and development of multi-dialect automatic speech recognition for Arabic. Deep neural networks are becoming an effective tool to solve sequential data problems, particularly, adopting an end-to-end training of…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-30 Abbas Raza Ali

Language comprehension and commonsense knowledge validation by machines are challenging tasks that are still under researched and evaluated for Arabic text. In this paper, we present a benchmark Arabic dataset for commonsense explanation.…

Computation and Language · Computer Science 2020-12-21 Saja AL-Tawalbeh , Mohammad AL-Smadi

Speech-based AI educational applications have gained significant interest in recent years, particularly for children. However, children speech research remains limited due to the lack of publicly available datasets, especially for…

Computation and Language · Computer Science 2026-03-24 Abdul Aziz Snoubara , Baraa Al_Maradni , Haya Al_Naal , Malek Al_Madrmani , Roaa Jdini , Seedra Zarzour , Khloud Al Jallad

The focus of language model evaluation has transitioned towards reasoning and knowledge-intensive tasks, driven by advancements in pretraining large models. While state-of-the-art models are partially trained on large Arabic texts,…

Although rare diseases are characterized by low prevalence, approximately 300 million people are affected by a rare disease. The early and accurate diagnosis of these conditions is a major challenge for general practitioners, who do not…

Computation and Language · Computer Science 2021-11-12 Isabel Segura-Bedmar , David Camino-Perdonas , Sara Guerrero-Aspizua

Instruction tuning has emerged as a prominent methodology for teaching Large Language Models (LLMs) to follow instructions. However, current instruction datasets predominantly cater to English or are derived from English-dominated LLMs,…

The largest dataset of Arabic speech mispronunciation detections in Egyptian dialogues is introduced. The dataset is composed of annotated audio files representing the top 100 words that are most frequently used in the Arabic language,…

Computation and Language · Computer Science 2021-11-03 Salah A. Aly , Abdelrahman Salah , Hesham M. Eraqi

In the real world, medical datasets often exhibit a long-tailed data distribution (i.e., a few classes occupy the majority of the data, while most classes have only a limited number of samples), which results in a challenging long-tailed…

Computer Vision and Pattern Recognition · Computer Science 2023-03-22 Lie Ju , Zhen Yu , Lin Wang , Xin Zhao , Xin Wang , Paul Bonnington , Zongyuan Ge