English
Related papers

Related papers: The SAMER Arabic Text Simplification Corpus

200 papers

Access to higher education is critical for minority populations and emergent bilingual students. However, the language used by higher education institutions to communicate with prospective students is often too complex; concretely, many…

Computation and Language · Computer Science 2022-09-13 Zachary W. Taylor , Maximus H. Chu , Junyi Jessy Li

In this paper, we present an ongoing effort in lexical semantic analysis and annotation of Modern Standard Arabic (MSA) text, a semi automatic annotation tool concerned with the morphologic, syntactic, and semantic levels of description.

Computation and Language · Computer Science 2016-05-25 Abdelaziz Lakhfif , Mohammed T. Laskri , Eric Atwell

Recently, with the rapid development in the fields of technology and the increasing amount of text t available on the internet, it has become urgent to develop effective tools for processing and understanding texts in a way that summaries…

Computation and Language · Computer Science 2024-06-13 Sari Masri , Yaqeen Raddad , Fidaa Khandaqji , Huthaifa I. Ashqar , Mohammed Elhenawy

Being able to understand information is a key factor for a self-determined life and society. It is also very important for participating in democratic processes. The study of automatic text simplification is often limited by the…

Computation and Language · Computer Science 2026-03-17 Stefan Bott , Verena Riegler , Horacio Saggion , Almudena Rascón Alcaina , Nouran Khallaf

In this paper, we present the annotation pipeline and the guidelines we wrote as part of an effort to create a large manually annotated Arabic author profiling dataset from various social media sources covering 16 Arabic countries and 11…

Computation and Language · Computer Science 2018-08-24 Wajdi Zaghouani , Anis Charfi

Abstractive text summarization is one of the areas influenced by the emergence of pre-trained language models. Current pre-training works in abstractive summarization give more points to the summaries with more words in common with the main…

Computation and Language · Computer Science 2021-09-10 Alireza Salemi , Emad Kebriaei , Ghazal Neisi Minaei , Azadeh Shakery

Data annotation is an important but time-consuming and costly procedure. To sort a text into two classes, the very first thing we need is a good annotation guideline, establishing what is required to qualify for each class. In the…

Computation and Language · Computer Science 2018-08-16 Imane Guellil , Ahsan Adeel , Faical Azouaou , Amir Hussain

With the expanding growth of Arabic electronic data on the web, extracting information, which is actually one of the major challenges of the question-answering, is essentially used for building corpus of documents. In fact, building a…

Information Retrieval · Computer Science 2018-05-24 Patrice Bellot , Wided Bakari , Mahmoud Neji

Arabic Optical Character Recognition (OCR) is essential for converting vast amounts of Arabic print media into digital formats. However, training modern OCR models, especially powerful vision-language models, is hampered by the lack of…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Omer Nacar , Yasser Al-Habashi , Serry Sibaee , Adel Ammar , Wadii Boulila

Text summarization has been intensively studied in many languages, and some languages have reached advanced stages. Yet, Arabic Text Summarization (ATS) is still in its developing stages. Existing ATS datasets are either small or lack…

Computation and Language · Computer Science 2022-10-26 Abdulaziz Alhamadani , Xuchao Zhang , Jianfeng He , Chang-Tien Lu

In this paper, we present a Modern Standard Arabic (MSA) Sentence difficulty classifier, which predicts the difficulty of sentences for language learners using either the CEFR proficiency levels or the binary classification as simple or…

Computation and Language · Computer Science 2021-03-09 Nouran Khallaf , Serge Sharoff

Sentiment Analysis (SA) is a major field of study in natural language processing, computational linguistics and information retrieval. Interest in SA has been constantly growing in both academia and industry over the recent years. Moreover,…

Computation and Language · Computer Science 2021-01-05 Pedram Hosseini , Ali Ahmadian Ramaki , Hassan Maleki , Mansoureh Anvari , Seyed Abolghasem Mirroshandel

International library standards require cataloguers to tediously input Romanization of their catalogue records for the benefit of library users without specific language expertise. In this paper, we present the first reported results on the…

Computation and Language · Computer Science 2021-03-15 Eryani Fadhl , Habash Nizar

Sentiment analysis (SA) has been, and is still, a thriving research area. However, the task of Arabic sentiment analysis (ASA) is still underrepresented in the body of research. This study offers the first in-depth and in-breadth analysis…

Computation and Language · Computer Science 2024-03-05 Latifah Almurqren , Ryan Hodgson , Alexandra Cristea

This demo paper presents a Google Docs add-on for automatic Arabic word-level readability visualization. The add-on includes a lemmatization component that is connected to a five-level readability lexicon and Arabic WordNet-based…

Computation and Language · Computer Science 2022-10-20 Reem Hazim , Hind Saddiki , Bashar Alhafni , Muhamed Al Khalil , Nizar Habash

Standing at the forefront of knowledge dissemination, digital libraries curate vast collections of scientific literature. However, these scholarly writings are often laden with jargon and tailored for domain experts rather than the general…

Computation and Language · Computer Science 2024-08-08 Haining Wang , Jason Clark

This paper introduces a pioneering English-Azerbaijani (Arabic Script) parallel corpus, designed to bridge the technological gap in language learning and machine translation (MT) for under-resourced languages. Consisting of 548,000 parallel…

This work consists of creating a system of the Computer Assisted Language Learning (CALL) based on a system of Automatic Speech Recognition (ASR) for the Arabic language using the tool CMU Sphinx3 [1], based on the approach of HMM. To this…

Computation and Language · Computer Science 2012-05-16 Naim Terbeh , Mounir Zrigui

We introduce the largest transcribed Arabic speech corpus, QASR, collected from the broadcast domain. This multi-dialect speech dataset contains 2,000 hours of speech sampled at 16kHz crawled from Aljazeera news channel. The dataset is…

Computation and Language · Computer Science 2021-06-25 Hamdy Mubarak , Amir Hussein , Shammur Absar Chowdhury , Ahmed Ali

Cross-Lingual SynthDocs is a large-scale synthetic corpus designed to address the scarcity of Arabic resources for Optical Character Recognition (OCR) and Document Understanding (DU). The dataset comprises over 2.5 million of samples,…

Computation and Language · Computer Science 2025-11-10 Haneen Al-Homoud , Asma Ibrahim , Murtadha Al-Jubran , Fahad Al-Otaibi , Yazeed Al-Harbi , Daulet Toibazar , Kesen Wang , Pedro J. Moreno