English
Related papers

Related papers: Creating a Large Multi-Layered Representational Re…

200 papers

In this paper, we introduce the first phase of a new dataset for offline Arabic handwriting recognition. The aim is to collect a very large dataset of isolated Arabic words that covers all letters of the alphabet in all possible shapes…

Computer Vision and Pattern Recognition · Computer Science 2014-11-19 Mohamed E. Hussein , Marwan Torki , Ahmed Elsallamy , Mahmoud Fayyaz

Cross-Lingual SynthDocs is a large-scale synthetic corpus designed to address the scarcity of Arabic resources for Optical Character Recognition (OCR) and Document Understanding (DU). The dataset comprises over 2.5 million of samples,…

Computation and Language · Computer Science 2025-11-10 Haneen Al-Homoud , Asma Ibrahim , Murtadha Al-Jubran , Fahad Al-Otaibi , Yazeed Al-Harbi , Daulet Toibazar , Kesen Wang , Pedro J. Moreno

Data contamination undermines the validity of Large Language Model evaluation by enabling models to rely on memorized benchmark content rather than true generalization. While prior work has proposed contamination detection methods, these…

Computation and Language · Computer Science 2026-01-22 Chaymaa Abbas , Nour Shamaa , Mariette Awad

Arabic text recognition is a challenging task because of the cursive nature of Arabic writing system, its joint writing scheme, the large number of ligatures and many other challenges. Deep Learning DL models achieved significant progress…

Computer Vision and Pattern Recognition · Computer Science 2020-09-07 Mohammad Fasha , Bassam Hammo , Nadim Obeid , Jabir Widian

In this paper, we propose efficient and less resource-intensive strategies for parsing of code-mixed data. These strategies are not constrained by in-domain annotations, rather they leverage pre-existing monolingual annotated resources for…

Computation and Language · Computer Science 2017-04-03 Irshad Ahmad Bhat , Riyaz Ahmad Bhat , Manish Shrivastava , Dipti Misra Sharma

ArzEn-MultiGenre is a parallel dataset of Egyptian Arabic song lyrics, novels, and TV show subtitles that are manually translated and aligned with their English counterparts. The dataset contains 25,557 segment pairs that can be used to…

Computation and Language · Computer Science 2025-08-05 Rania Al-Sabbagh

In contrast to many decades of research on oral code-switching, the study of written multilingual productions has only recently enjoyed a surge of interest. Many open questions remain regarding the sociolinguistic underpinnings of written…

Computation and Language · Computer Science 2019-09-02 Ella Rabinovich , Masih Sultani , Suzanne Stevenson

In NLP, text classification is one of the primary problems we try to solve and its uses in language analyses are indisputable. The lack of labeled training data made it harder to do these tasks in low resource languages like Amharic. The…

Computation and Language · Computer Science 2021-03-11 Israel Abebe Azime , Nebil Mohammed

The rise of social media such as blogs and social networks has fueled interest in sentiment analysis. With the proliferation of reviews, ratings, recommendations and other forms of online expression, online opinion has turned into a kind of…

Computation and Language · Computer Science 2015-05-13 Hossam S. Ibrahim , Sherif M. Abdou , Mervat Gheith

Code-switching -- the natural alternation between two languages within a single utterance -- remains one of the most challenging and under-studied conditions for automatic speech recognition (ASR). We present a benchmark evaluating five…

Computation and Language · Computer Science 2026-05-25 Sajjad Abdoli , Ghassan Al-Sumaidaee , Clayton W. Taylor , Ahmad ElShiekh , Ahmed Rashad

Open Arabic large language models split into two classes: sub-1B multilingual models that treat Arabic as an afterthought (Qwen2.5-0.5B, Falcon-H1-0.5B), and 7B-70B Arabic-specialized models that require a server to run (Jais, AceGPT,…

Computation and Language · Computer Science 2026-05-29 Jaber Jaber , Osama Jaber

Despite the advances in neural text to speech (TTS), many Arabic dialectal varieties remain marginally addressed, with most resources concentrated on Modern Spoken Arabic (MSA) and Gulf dialects, leaving Egyptian Arabic -- the most widely…

Computation and Language · Computer Science 2026-03-30 Ahmed Khaled Khamis , Hesham Ali

Arabic dialect identification (ADI) systems are essential for large-scale data collection pipelines that enable the development of inclusive speech technologies for Arabic language varieties. However, the reliability of current ADI systems…

Computation and Language · Computer Science 2025-06-02 Badr M. Abdullah , Matthew Baas , Bernd Möbius , Dietrich Klakow

In this paper, we address the significant gap in Arabic natural language processing (NLP) resources by introducing ArabicaQA, the first large-scale dataset for machine reading comprehension and open-domain question answering in Arabic. This…

Computation and Language · Computer Science 2024-03-27 Abdelrahman Abdallah , Mahmoud Kasem , Mahmoud Abdalla , Mohamed Mahmoud , Mohamed Elkasaby , Yasser Elbendary , Adam Jatowt

This paper describes AraS2P, our speech-to-phonemes system submitted to the Iqra'Eval 2025 Shared Task. We adapted Wav2Vec2-BERT via Two-Stage training strategy. In the first stage, task-adaptive continue pretraining was performed on…

Computation and Language · Computer Science 2025-09-30 Bassam Matar , Mohamed Fayed , Ayman Khalafallah

Social media has become a crucial arena for shaping public narratives during armed conflicts, providing space for both harmful and constructive communication. While hate speech and misinformation have been widely studied, expressions that…

Computation and Language · Computer Science 2026-05-25 Esra'a Sharqawi , Wajdi Zaghouani

Along with the COVID-19 pandemic, an "infodemic" of false and misleading information has emerged and has complicated the COVID-19 response efforts. Social networking sites such as Facebook and Twitter have contributed largely to the spread…

Computation and Language · Computer Science 2021-05-10 Mohamed Seghir Hadj Ameur , Hassina Aliane

Handwritten character recognition has been the center of research and a benchmark problem in the sector of pattern recognition and artificial intelligence, and it continues to be a challenging research topic. Due to its enormous application…

Computer Vision and Pattern Recognition · Computer Science 2022-09-09 Akm Ashiquzzaman , Abdul Kawsar Tushar , Md Ashiqur Rahman

Building dialogues systems interaction has recently gained considerable attention, but most of the resources and systems built so far are tailored to English and other Indo-European languages. The need for designing systems for other…

Computation and Language · Computer Science 2015-05-13 AbdelRahim A. Elmadany , Sherif M. Abdou , Mervat Gheith

Question answering(QA) is one of the most challenging yet widely investigated problems in Natural Language Processing (NLP). Question-answering (QA) systems try to produce answers for given questions. These answers can be generated from…

Computation and Language · Computer Science 2025-08-06 Kholoud Alsubhi , Amani Jamal , Areej Alhothali