中文
相关论文

相关论文: SALMA: Arabic Sense-Annotated Corpus and WSD Bench…

200 篇论文

Arabic Sign Language (ArSL) and its dialects serve approximately 400 million Arabic speakers worldwide, yet the community lacks high-quality 3D parametric annotations and specialized reconstruction methods for avatar generation. We address…

计算机视觉与模式识别 · 计算机科学 2026-05-14 Eyad Alghamdi , Sattam Altuuaim , Obay Ghulam , Abdulrahman Qutah , Yousef Basoodan

In this paper, an approach for hate speech detection against women in Arabic community on social media (e.g. Youtube) is proposed. In the literature, similar works have been presented for other languages such as English. However, to the…

计算与语言 · 计算机科学 2021-04-06 Imane Guellil , Ahsan Adeel , Faical Azouaou , Mohamed Boubred , Yousra Houichi , Akram Abdelhaq Moumna

Question semantic similarity (Q2Q) is a challenging task that is very useful in many NLP applications, such as detecting duplicate questions and question answering systems. In this paper, we present the results and findings of the shared…

计算与语言 · 计算机科学 2019-09-24 Haitham Seelawi , Ahmad Mustafa , Hesham Al-Bataineh , Wael Farhan , Hussein T. Al-Natsheh

High-quality Machine Translation (MT) evaluation relies heavily on human judgments. Comprehensive error classification methods, such as Multidimensional Quality Metrics (MQM), are expensive as they are time-consuming and can only be done by…

This survey provides the first systematic review of Arabic LLM benchmarks, analyzing 40+ evaluation benchmarks across NLP tasks, knowledge domains, cultural understanding, and specialized capabilities. We propose a taxonomy organizing…

We introduce the Tarab Corpus, a large-scale cultural and linguistic resource that brings together Arabic song lyrics and poetry within a unified analytical framework. The corpus comprises 2.56 million verses and more than 13.5 million…

计算与语言 · 计算机科学 2026-03-18 Mo El-Haj

Event-argument extraction is a challenging task, particularly in Arabic due to sparse linguistic resources. To fill this gap, we introduce the \hadath corpus ($550$k tokens) as an extension of Wojood, enriched with event-argument…

计算与语言 · 计算机科学 2024-08-01 Alaa Aljabari , Lina Duaibes , Mustafa Jarrar , Mohammed Khalilia

Over the recent decades, there has been a significant increase and development of resources for Arabic natural language processing. This includes the task of exploring Arabic Language Sentiment Analysis (ALSA) from Arabic utterances in both…

计算与语言 · 计算机科学 2021-09-16 Azza Abugharsa

The complexities of Arabic language in morphology, orthography and dialects makes sentiment analysis for Arabic more challenging. Also, text feature extraction from short messages like tweets, in order to gauge the sentiment, makes this…

计算与语言 · 计算机科学 2018-10-17 Abdulaziz M. Alayba , Vasile Palade , Matthew England , Rahat Iqbal

This paper presents Wojood, a corpus for Arabic nested Named Entity Recognition (NER). Nested entities occur when one entity mention is embedded inside another entity mention. Wojood consists of about 550K Modern Standard Arabic (MSA) and…

计算与语言 · 计算机科学 2022-05-24 Mustafa Jarrar , Mohammed Khalilia , Sana Ghanem

We present a word-sense induction method based on pre-trained masked language models (MLMs), which can cheaply scale to large vocabularies and large corpora. The result is a corpus which is sense-tagged according to a corpus-derived sense…

计算与语言 · 计算机科学 2022-03-22 Matan Eyal , Shoval Sadde , Hillel Taub-Tabib , Yoav Goldberg

Arabic is a semitic language characterized by a complex and rich morphology. The exceptional degree of ambiguity in the writing system, the rich morphology, and the highly complex word formation process of roots and patterns all contribute…

计算机视觉与模式识别 · 计算机科学 2014-12-25 Ibrahim Abdelaziz , Sherif Abdou

This paper describes two systems that were used by the authors for addressing Arabic Sentiment Analysis as part of SemEval-2017, task 4. The authors participated in three Arabic related subtasks which are: Subtask A (Message Polarity…

计算与语言 · 计算机科学 2017-10-25 Samhaa R. El-Beltagy , Mona El Kalamawy , Abu Bakr Soliman

The presented work aims at generating a systematically annotated corpus that can support the enhancement of sentiment analysis tasks in Telugu using word-level sentiment annotations. From OntoSenseNet, we extracted 11,000 adjectives, 253…

计算与语言 · 计算机科学 2018-07-05 Sreekavitha Parupalli , Vijjini Anvesh Rao , Radhika Mamidi

Word Sense Disambiguation (WSD) has been widely evaluated using the semantic frameworks of WordNet, BabelNet, and the Oxford Dictionary of English. However, for the UCREL Semantic Analysis System (USAS) framework, no open extensive…

Many annotation tools have been developed, covering a wide variety of tasks and providing features like user management, pre-processing, and automatic labeling. However, all of these tools use Graphical User Interfaces, and often require…

计算与语言 · 计算机科学 2020-06-05 Jonathan K. Kummerfeld

This paper proposes a methodology to prepare corpora in Arabic language from online social network (OSN) and review site for Sentiment Analysis (SA) task. The paper also proposes a methodology for generating a stopword list from the…

计算与语言 · 计算机科学 2014-10-07 Walaa Medhat , Ahmed H. Yousef , Hoda Korashy

We introduce Konooz, a novel multi-dimensional corpus covering 16 Arabic dialects across 10 domains, resulting in 160 distinct corpora. The corpus comprises about 777k tokens, carefully collected and manually annotated with 21 entity types…

计算与语言 · 计算机科学 2025-06-17 Nagham Hamad , Mohammed Khalilia , Mustafa Jarrar

This paper introduces Ta'keed, an explainable Arabic automatic fact-checking system. While existing research often focuses on classifying claims as "True" or "False," there is a limited exploration of generating explanations for claim…

计算与语言 · 计算机科学 2024-01-26 Saud Althabiti , Mohammad Ammar Alsalka , Eric Atwell