English
Related papers

Related papers: 1.5 billion words Arabic Corpus

200 papers

One of the most major and essential tasks in natural language processing is machine translation that is now highly dependent upon multilingual parallel corpora. Through this paper, we introduce the biggest Persian-English parallel corpus…

Computation and Language · Computer Science 2020-02-03 Omid Kashefi

Developing robust automatic speech recognition (ASR) systems for Arabic requires effective strategies to manage its diversity. Existing ASR systems mainly cover the modern standard Arabic (MSA) variety and few high-resource dialects, but…

Computation and Language · Computer Science 2025-06-02 Amirbek Djanibekov , Hawau Olamide Toyin , Raghad Alshalan , Abdullah Alitr , Hanan Aldarmaki

Arabic calligraphy represents one of the richest visual traditions of the Arabic language, blending linguistic meaning with artistic form. Although multimodal models have advanced across languages, their ability to process Arabic script,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 Shubham Patle , Sara Ghaboura , Hania Tariq , Mohammad Usman Khan , Omkar Thawakar , Rao Muhammad Anwer , Salman Khan

SALMA, the first Arabic sense-annotated corpus, consists of ~34K tokens, which are all sense-annotated. The corpus is annotated using two different sense inventories simultaneously (Modern and Ghani). SALMA novelty lies in how tokens and…

Computation and Language · Computer Science 2023-10-31 Mustafa Jarrar , Sanad Malaysha , Tymaa Hammouda , Mohammed Khalilia

We present Multilingual Open Text (MOT), a new multilingual corpus containing text in 44 languages, many of which have limited existing text resources for natural language processing. The first release of the corpus contains over 2.8…

Computation and Language · Computer Science 2022-06-10 Chester Palen-Michel , June Kim , Constantine Lignos

We present the Manuscripts of Handwritten Arabic~(Muharaf) dataset, which is a machine learning dataset consisting of more than 1,600 historic handwritten page images transcribed by experts in archival Arabic. Each document image is…

Computer Vision and Pattern Recognition · Computer Science 2025-02-06 Mehreen Saeed , Adrian Chan , Anupam Mijar , Joseph Moukarzel , Georges Habchi , Carlos Younes , Amin Elias , Chau-Wai Wong , Akram Khater

We present the Knesset Corpus, a corpus of Hebrew parliamentary proceedings containing over 30 million sentences (over 384 million tokens) from all the (plenary and committee) protocols held in the Israeli parliament between 1998 and 2022.…

Computation and Language · Computer Science 2025-06-02 Gili Goldin , Nick Howell , Noam Ordan , Ella Rabinovich , Shuly Wintner

Language recognition has been significantly advanced in recent years by means of modern machine learning methods such as deep learning and benchmarks with rich annotations. However, research is still limited in low-resource formal…

Computation and Language · Computer Science 2020-06-03 Hadi Abdi Khojasteh , Ebrahim Ansari , Mahdi Bohlouli

We present AraLingBench: a fully human annotated benchmark for evaluating the Arabic linguistic competence of large language models (LLMs). The benchmark spans five core categories: grammar, morphology, spelling, reading comprehension, and…

Computation and Language · Computer Science 2025-12-10 Mohammad Zbeeb , Hasan Abed Al Kader Hammoud , Sina Mukalled , Nadine Rizk , Fatima Karnib , Issam Lakkis , Ammar Mohanna , Bernard Ghanem

This survey provides the first systematic review of Arabic LLM benchmarks, analyzing 40+ evaluation benchmarks across NLP tasks, knowledge domains, cultural understanding, and specialized capabilities. We propose a taxonomy organizing…

Many European languages possess rich biblical translation histories, yet existing corpora - in prioritizing linguistic breadth - often fail to capture this depth. To address this gap, we introduce a multilingual corpus of 651 New Testament…

Computation and Language · Computer Science 2026-05-14 Maciej Rapacz , Aleksander Smywiński-Pohl

The debut of chatGPT and BARD has popularized instruction following text generation using LLMs, where a user can interrogate an LLM using natural language requests and obtain natural language answers that matches their requests. Training…

Computation and Language · Computer Science 2024-08-13 Abdelrahman El-Sheikh , Ahmed Elmogtaba , Kareem Darwish , Muhammad Elmallah , Ashraf Elneima , Hassan Sawaf

Recognizing causal elements and causal relations in text is one of the challenging issues in natural language processing; specifically, in low resource languages such as Persian. In this research we prepare a causality human annotated…

Computation and Language · Computer Science 2021-06-29 Zeinab Rahimi , Mehrnoush ShamsFard

In this work we present our expert system of Automatic reading or speech synthesis based on a text written in Standard Arabic, our work is carried out in two great stages: the creation of the sound data base, and the transformation of the…

Computation and Language · Computer Science 2014-05-09 Tebbi Hanane , Azzoune Hamid

SinhaLegal introduces a Sinhala legislative text corpus containing approximately 2 million words across 1,206 legal documents. The dataset includes two types of legal documents: 1,065 Acts dated from 1981 to 2014 and 141 Bills from 2010 to…

Computation and Language · Computer Science 2026-03-06 Minduli Lasandi , Nevidu Jayatilleke

The algorithm of the creation texts parallel corpora was presented. The algorithm is based on the use of "key words" in text documents, and on the means of their automated translation. Key words were singled out by means of using Russian…

Computation and Language · Computer Science 2008-07-03 D. V. Lande , V. V. Zhygalo

Arabic handwriting is a consonantal and cursive writing. The analysis of Arabic script is further complicated due to obligatory dots/strokes that are placed above or below most letters and usually written delayed in order. Due to…

Computer Vision and Pattern Recognition · Computer Science 2015-10-20 Ibrahim Abdelaziz , Sherif Abdou , Hassanin Al-Barhamtoshy

We release a corpus of 43 million atomic edits across 8 languages. These edits are mined from Wikipedia edit history and consist of instances in which a human editor has inserted a single contiguous phrase into, or deleted a single…

Computation and Language · Computer Science 2018-08-29 Manaal Faruqui , Ellie Pavlick , Ian Tenney , Dipanjan Das

In the context of arabic Information Retrieval Systems (IRS) guided by arabic ontology and to enable those systems to better respond to user requirements, this paper aims to representing documents and queries by the best concepts extracted…

Information Retrieval · Computer Science 2013-06-28 Mohammed Alaeddine Abderrahim , Mohammed El Amine Abderrahim , Mohammed Amine Chikh

In this paper, we introduce SaudiBERT, a monodialect Arabic language model pretrained exclusively on Saudi dialectal text. To demonstrate the model's effectiveness, we compared SaudiBERT with six different multidialect Arabic language…

Computation and Language · Computer Science 2024-05-13 Faisal Qarah