English
Related papers

Related papers: An Enhanced Corpus for Arabic Newspapers Comments

200 papers

The rapid advancement of social media enables us to analyze user opinions. In recent times, sentiment analysis has shown a prominent research gap in understanding human sentiment based on the content shared on social media. Although…

Computation and Language · Computer Science 2024-03-12 Md Arid Hasan

The main aim of this study is the assessment and discussion of a model for hand-written Arabic through segmentation. The framework is proposed based on three steps: pre-processing, segmentation, and evaluation. In the pre-processing step,…

Computer Vision and Pattern Recognition · Computer Science 2021-01-11 Nisreen AbdAllah , Serestina Viriri

In this research, we continuously collect data from the RSS feeds of traditional news sources. We apply several pre-trained implementations of named entity recognition (NER) tools, quantifying the success of each implementation. We also…

Computers and Society · Computer Science 2020-06-11 Ashwini Badgujar , Sheng Chen , Andrew Wang , Kai Yu , Paul Intrevado , David Guy Brizan

Representation of semantic information contained in the words is needed for any Arabic Text Mining applications. More precisely, the purpose is to better take into account the semantic dependencies between words expressed by the…

Computation and Language · Computer Science 2012-12-18 Hanane Froud , Abdelmonaim Lachkar , Said Alaoui Ouatik

This paper presents an attempt to build a Modern Standard Arabic (MSA) sentence-level simplification system. We experimented with sentence simplification using two approaches: (i) a classification approach leading to lexical simplification…

Computation and Language · Computer Science 2022-04-21 Nouran Khallaf , Serge Sharoff

The first step of processing a question in Question Answering(QA) Systems is to carry out a detailed analysis of the question for the purpose of determining what it is asking for and how to perfectly approach answering it. Our Question…

Computation and Language · Computer Science 2017-01-12 Waheeb Ahmed , Dr. Anto P Babu

Fake news and deceptive machine-generated text are serious problems threatening modern societies, including in the Arab world. This motivates work on detecting false and manipulated stories online. However, a bottleneck for this research is…

Computation and Language · Computer Science 2020-11-09 El Moatez Billah Nagoudi , AbdelRahim Elmadany , Muhammad Abdul-Mageed , Tariq Alhindi , Hasan Cavusoglu

Comparable texts are topic-aligned documents in multiple languages that are not direct translations. They are valuable for understanding how a topic is discussed across languages. This research studies differences in sentiments and emotions…

Computation and Language · Computer Science 2025-08-06 Motaz Saad , David Langlois , Kamel Smaili

This paper describes a novel study on using `Attention Mask' input in transformers and using this approach for detecting offensive content in both English and Persian languages. The paper's principal focus is to suggest a methodology to…

Computation and Language · Computer Science 2021-10-12 Peyman Alavi , Pouria Nikvand , Mehrnoush Shamsfard

We present ARETA, an automatic error type annotation system for Modern Standard Arabic. We design ARETA to address Arabic's morphological richness and orthographic ambiguity. We base our error taxonomy on the Arabic Learner Corpus (ALC)…

Computation and Language · Computer Science 2021-09-17 Riadh Belkebir , Nizar Habash

Many AI detection models have been developed to counter the presence of articles created by artificial intelligence (AI). However, if a human-authored article is slightly polished by AI, a shift will occur in the borderline decision of…

Computation and Language · Computer Science 2025-12-03 Saleh Almohaimeed , Saad Almohaimeed , Mousa Jari , Khaled A. Alobaid , Fahad Alotaibi

Framing detection in Arabic social media is difficult due to interpretive ambiguity, cultural grounding, and limited reliable supervision. Existing LLM-based weak supervision methods typically rely on label aggregation, which is brittle…

Computation and Language · Computer Science 2026-03-06 Rabab Alkhalifa

Lack of available resources such as text corpora for low-resource languages seriously hinders research on natural language processing and computational linguistics. This paper presents AlbMoRe, a corpus of 800 sentiment annotated movie…

Computation and Language · Computer Science 2023-06-16 Erion Çano

Given the number of Arabic speakers worldwide and the notably large amount of content in the web today in some fields such as law, medicine, or even news, documents of considerable length are produced regularly. Classifying those documents…

Computation and Language · Computer Science 2023-05-08 Muhammad AL-Qurishi

Arabic Documents Clustering is an important task for obtaining good results with the traditional Information Retrieval (IR) systems especially with the rapid growth of the number of online documents present in Arabic language. Documents…

Information Retrieval · Computer Science 2013-02-08 Hanane Froud , Abdelmonaime Lachkar , Said Alaoui Ouatik

Hybrid approaches for automatic vowelization of Arabic texts are presented in this article. The process is made up of two modules. In the first one, a morphological analysis of the text words is performed using the open source morphological…

Computation and Language · Computer Science 2014-10-13 Mohamed Bebah , Chennoufi Amine , Mazroui Azzeddine , Lakhouaja Abdelhak

Identifying hate speech content in the Arabic language is challenging due to the rich quality of dialectal variations. This study introduces a multilabel hate speech dataset in the Arabic language. We have collected 10000 Arabic tweets and…

Computation and Language · Computer Science 2025-05-26 Wajdi Zaghouani , Md. Rafiul Biswas

This paper investigates the optimization of propaganda technique detection in Arabic text, including tweets \& news paragraphs, from ArAIEval shared task 1. Our approach involves fine-tuning the AraBERT v2 model with a neural network…

Computation and Language · Computer Science 2024-07-02 Abrar Abir , Kemal Oflazer

Arabic dialects form a diverse continuum, yet NLP models often treat them as discrete categories. Recent work addresses this issue by modeling dialectness as a continuous variable, notably through the Arabic Level of Dialectness (ALDi).…

Computation and Language · Computer Science 2025-08-26 Sanad Shaban , Nizar Habash

Unsupervised text classification, with its most common form being sentiment analysis, used to be performed by counting words in a text that were stored in a lexicon, which assigns each word to one class or as a neutral word. In recent…

Computation and Language · Computer Science 2025-06-26 Kai-Robin Lange , Jonas Rieger , Carsten Jentsch