中文
相关论文

相关论文: A Large Scale Corpus of Gulf Arabic

200 篇论文

This study is an attempt to build a contemporary linguistic corpus for Arabic language. The corpus produced, is a text corpus includes more than five million newspaper articles. It contains over a billion and a half words in total, out of…

计算与语言 · 计算机科学 2016-11-15 Ibrahim Abu El-khair

Arabic is a widely-spoken language with a rich and long history spanning more than fourteen centuries. Yet existing Arabic corpora largely focus on the modern period or lack sufficient diachronic information. We develop a large-scale,…

计算与语言 · 计算机科学 2016-12-30 Yonatan Belinkov , Alexander Magidow , Maxim Romanov , Avi Shmidman , Moshe Koppel

We present the SAMER Corpus, the first manually annotated Arabic parallel corpus for text simplification targeting school-aged learners. Our corpus comprises texts of 159K words selected from 15 publicly available Arabic fiction novels most…

计算与语言 · 计算机科学 2024-04-30 Bashar Alhafni , Reem Hazim , Juan Piñeros Liberato , Muhamed Al Khalil , Nizar Habash

Arabic is a widely-spoken language with a long and rich history, but existing corpora and language technology focus mostly on modern Arabic and its varieties. Therefore, studying the history of the language has so far been mostly limited to…

计算与语言 · 计算机科学 2018-09-12 Yonatan Belinkov , Alexander Magidow , Alberto Barrón-Cedeño , Avi Shmidman , Maxim Romanov

We present the QuranMorph corpus, a morphologically annotated corpus for the Quran (77,429 tokens). Each token in the QuranMorph was manually lemmatized and tagged with its part-of-speech by three expert linguists. The lemmatization process…

计算与语言 · 计算机科学 2025-06-24 Diyam Akra , Tymaa Hammouda , Mustafa Jarrar

We introduce the Tarab Corpus, a large-scale cultural and linguistic resource that brings together Arabic song lyrics and poetry within a unified analytical framework. The corpus comprises 2.56 million verses and more than 13.5 million…

计算与语言 · 计算机科学 2026-03-18 Mo El-Haj

This paper introduces the Balanced Arabic Readability Evaluation Corpus (BAREC), a large-scale, fine-grained dataset for Arabic readability assessment. BAREC consists of 69,441 sentences spanning 1+ million words, carefully curated to cover…

计算与语言 · 计算机科学 2025-06-17 Khalid N. Elmadani , Nizar Habash , Hanada Taha-Thomure

In this paper, we present the annotation pipeline and the guidelines we wrote as part of an effort to create a large manually annotated Arabic author profiling dataset from various social media sources covering 16 Arabic countries and 11…

计算与语言 · 计算机科学 2018-08-24 Wajdi Zaghouani , Anis Charfi

In this paper, we present an ongoing effort in lexical semantic analysis and annotation of Modern Standard Arabic (MSA) text, a semi automatic annotation tool concerned with the morphologic, syntactic, and semantic levels of description.

计算与语言 · 计算机科学 2016-05-25 Abdelaziz Lakhfif , Mohammed T. Laskri , Eric Atwell

In this paper, we present Arap-Tweet, which is a large-scale and multi-dialectal corpus of Tweets from 11 regions and 16 countries in the Arab world representing the major Arabic dialectal varieties. To build this corpus, we collected data…

计算与语言 · 计算机科学 2018-08-24 Wajdi Zaghouani , Anis Charfi

The availability of corpora is a major factor in building natural language processing applications. However, the costs of acquiring corpora can prevent some researchers from going further in their endeavours. The ease of access to freely…

计算与语言 · 计算机科学 2017-02-28 Wajdi Zaghouani

Language models (LMs) have introduced a major paradigm shift in Natural Language Processing (NLP) modeling where large pre-trained LMs became integral to most of the NLP tasks. The LMs are intelligent enough to find useful and relevant…

计算与语言 · 计算机科学 2023-05-09 Abbas Raza Ali , Muhammad Ajmal Siddiqui , Rema Algunaibet , Hasan Raza Ali

The processing of the Arabic language is a complex field of research. This is due to many factors, including the complex and rich morphology of Arabic, its high degree of ambiguity, and the presence of several regional varieties that need…

计算与语言 · 计算机科学 2022-05-20 Karim El Haff , Mustafa Jarrar , Tymaa Hammouda , Fadi Zaraket

This paper describes a web-based corpus of global language use with a focus on how this corpus can be used for data-driven language mapping. First, the corpus provides a representation of where national varieties of major languages are used…

计算与语言 · 计算机科学 2020-04-03 Jonathan Dunn

Automatic readability assessment is relevant to building NLP applications for education, content analysis, and accessibility. However, Arabic readability assessment is a challenging task due to Arabic's morphological richness and limited…

计算与语言 · 计算机科学 2024-07-04 Juan Piñeros Liberato , Bashar Alhafni , Muhamed Al Khalil , Nizar Habash

Arabic is recognised as the 4th most used language of the Internet. Arabic has three main varieties: (1) classical Arabic (CA), (2) Modern Standard Arabic (MSA), (3) Arabic Dialect (AD). MSA and AD could be written either in Arabic or in…

计算与语言 · 计算机科学 2019-03-08 Imane Guellil , Houda Saâdane , Faical Azouaou , Billel Gueni , Damien Nouvel

The rise of large language models (LLMs) has transformed numerous natural language processing (NLP) tasks, yet their performance in low and mid-resource languages, such as Farsi, still lags behind resource-rich languages like English. To…

计算与语言 · 计算机科学 2024-12-24 Sadra Sabouri , Elnaz Rahmati , Soroush Gooran , Hossein Sameti

There are numerous complex and rich morphological features in the Arabic language, which are highly useful when analyzing traditional Arabic textbooks, especially in the literary and religious contexts, and help in understanding the meaning…

计算与语言 · 计算机科学 2025-01-24 Huda AlShuhayeb , Behrouz Minaei-Bidgoli , Mohammad E. Shenassa , Sayyed-Ali Hossayni

In recent years, Large Language Models have revolutionized the field of natural language processing, showcasing an impressive rise predominantly in English-centric domains. These advancements have set a global benchmark, inspiring…

计算与语言 · 计算机科学 2024-05-06 Manel Aloui , Hasna Chouikhi , Ghaith Chaabane , Haithem Kchaou , Chehir Dhaouadi

This paper presents Nabra, a corpora of Syrian Arabic dialects with morphological annotations. A team of Syrian natives collected more than 6K sentences containing about 60K words from several sources including social media posts, scripts…

计算与语言 · 计算机科学 2023-10-27 Amal Nayouf , Tymaa Hammouda , Mustafa Jarrar , Fadi Zaraket , Mohamad-Bassam Kurdy
‹ 上一页 1 2 3 10 下一页 ›