中文
相关论文

相关论文: A Large and Balanced Corpus for Fine-grained Arabi…

200 篇论文

High-quality WordNets are crucial for achieving high-quality results in NLP applications that rely on such resources. However, the wordnets of most languages suffer from serious issues of correctness and completeness with respect to the…

计算与语言 · 计算机科学 2024-04-01 Abed Alhakim Freihat , Hadi Khalilia , Gábor Bella , Fausto Giunchiglia

We propose a new method for evaluating the readability of simplified sentences through pair-wise ranking. The validity of the method is established through in-corpus and cross-corpus evaluation experiments. The approach correctly identifies…

计算与语言 · 计算机科学 2016-03-22 Sowmya Vajjala , Detmar Meurers

Spelling correction is the task of identifying spelling mistakes, typos, and grammatical mistakes in a given text and correcting them according to their context and grammatical structure. This work introduces "AraSpell," a framework for…

计算与语言 · 计算机科学 2024-05-14 Mahmoud Salhab , Faisal Abu-Khzam

We describe an Arabic-Hebrew parallel corpus of TED talks built upon WIT3, the Web inventory that repurposes the original content of the TED website in a way which is more convenient for MT researchers. The benchmark consists of about 2,000…

计算与语言 · 计算机科学 2016-10-04 Mauro Cettolo

High-quality parallel corpora are essential for Machine Translation (MT) research and translation teaching. However, Arabic-English resources remain scarce and existing datasets mainly consist of simple one-to-one mappings. In this paper,…

计算与语言 · 计算机科学 2026-01-05 Baorong Huang , Ali Asiri

We introduce ALARB, a dataset and suite of tasks designed to evaluate the reasoning capabilities of large language models (LLMs) within the Arabic legal domain. While existing Arabic benchmarks cover some knowledge-intensive tasks such as…

We introduce ArVoice, a multi-speaker Modern Standard Arabic (MSA) speech corpus with diacritized transcriptions, intended for multi-speaker speech synthesis, and can be useful for other tasks such as speech-based diacritic restoration,…

计算与语言 · 计算机科学 2025-05-28 Hawau Olamide Toyin , Rufael Marew , Humaid Alblooshi , Samar M. Magdy , Hanan Aldarmaki

We present ARETA, an automatic error type annotation system for Modern Standard Arabic. We design ARETA to address Arabic's morphological richness and orthographic ambiguity. We base our error taxonomy on the Arabic Learner Corpus (ALC)…

计算与语言 · 计算机科学 2021-09-17 Riadh Belkebir , Nizar Habash

In this paper, we present a Modern Standard Arabic (MSA) Sentence difficulty classifier, which predicts the difficulty of sentences for language learners using either the CEFR proficiency levels or the binary classification as simple or…

计算与语言 · 计算机科学 2021-03-09 Nouran Khallaf , Serge Sharoff

Despite its significance, Arabic, a linguistically rich and morphologically complex language, faces the challenge of being under-resourced. The scarcity of large annotated datasets hampers the development of accurate tools for subjectivity…

计算与语言 · 计算机科学 2026-03-02 Slimane Bellaouar , Attia Nehar , Soumia Souffi , Mounia Bouameur

The Arabic language has undergone notable transformations over time, including the emergence of new vocabulary, the obsolescence of others, and shifts in word usage. This evolution is evident in the distinction between the classical and…

计算与语言 · 计算机科学 2026-01-23 Zainab Alhathloul , Irfan Ahmad

SALMA, the first Arabic sense-annotated corpus, consists of ~34K tokens, which are all sense-annotated. The corpus is annotated using two different sense inventories simultaneously (Modern and Ghani). SALMA novelty lies in how tokens and…

计算与语言 · 计算机科学 2023-10-31 Mustafa Jarrar , Sanad Malaysha , Tymaa Hammouda , Mohammed Khalilia

Sentiment analysis (SA) has been, and is still, a thriving research area. However, the task of Arabic sentiment analysis (ASA) is still underrepresented in the body of research. This study offers the first in-depth and in-breadth analysis…

计算与语言 · 计算机科学 2024-03-05 Latifah Almurqren , Ryan Hodgson , Alexandra Cristea

The processing of the Arabic language is a complex field of research. This is due to many factors, including the complex and rich morphology of Arabic, its high degree of ambiguity, and the presence of several regional varieties that need…

计算与语言 · 计算机科学 2022-05-20 Karim El Haff , Mustafa Jarrar , Tymaa Hammouda , Fadi Zaraket

Language models (LMs) have introduced a major paradigm shift in Natural Language Processing (NLP) modeling where large pre-trained LMs became integral to most of the NLP tasks. The LMs are intelligent enough to find useful and relevant…

计算与语言 · 计算机科学 2023-05-09 Abbas Raza Ali , Muhammad Ajmal Siddiqui , Rema Algunaibet , Hasan Raza Ali

Arabic is a linguistically and culturally rich language with a vast vocabulary that spans scientific, religious, and literary domains. Yet, large-scale lexical datasets linking Arabic words to precise definitions remain limited. We present…

计算与语言 · 计算机科学 2026-01-30 Serry Sibaee , Yasser Alhabashi , Nadia Sibai , Yara Farouk , Adel Ammar , Sawsan AlHalawani , Wadii Boulila

Grammatical error correction (GEC) is a well-explored problem in English with many existing models and datasets. However, research on GEC in morphologically rich languages has been limited due to challenges such as data scarcity and…

计算与语言 · 计算机科学 2023-11-10 Bashar Alhafni , Go Inoue , Christian Khairallah , Nizar Habash

An automated approach to text readability assessment is essential to a language and can be a powerful tool for improving the understandability of texts written and published in that language. However, the Persian language, which is spoken…

计算与语言 · 计算机科学 2020-04-23 Hamid Mohammadi , Seyed Hossein Khasteh

Named Entity Recognition is an information extraction task that serves as a preprocessing step for other natural language processing tasks, such as machine translation, information retrieval, and question answering. Named entity recognition…

计算与语言 · 计算机科学 2022-07-05 Ebrahim Chekol Jibril , A. Cüneyd Tantğ

We introduce the largest transcribed Arabic speech corpus, QASR, collected from the broadcast domain. This multi-dialect speech dataset contains 2,000 hours of speech sampled at 16kHz crawled from Aljazeera news channel. The dataset is…

计算与语言 · 计算机科学 2021-06-25 Hamdy Mubarak , Amir Hussein , Shammur Absar Chowdhury , Ahmed Ali