中文
相关论文

相关论文: 1.5 billion words Arabic Corpus

200 篇论文

Arabic is a widely-spoken language with a rich and long history spanning more than fourteen centuries. Yet existing Arabic corpora largely focus on the modern period or lack sufficient diachronic information. We develop a large-scale,…

计算与语言 · 计算机科学 2016-12-30 Yonatan Belinkov , Alexander Magidow , Maxim Romanov , Avi Shmidman , Moshe Koppel

Arabic is a widely-spoken language with a long and rich history, but existing corpora and language technology focus mostly on modern Arabic and its varieties. Therefore, studying the history of the language has so far been mostly limited to…

计算与语言 · 计算机科学 2018-09-12 Yonatan Belinkov , Alexander Magidow , Alberto Barrón-Cedeño , Avi Shmidman , Maxim Romanov

Most Arabic natural language processing tools and resources are developed to serve Modern Standard Arabic (MSA), which is the official written language in the Arab World. Some Dialectal Arabic varieties, notably Egyptian Arabic, have…

计算与语言 · 计算机科学 2016-09-13 Salam Khalifa , Nizar Habash , Dana Abdulrahim , Sara Hassan

This paper introduces the Balanced Arabic Readability Evaluation Corpus (BAREC), a large-scale, fine-grained dataset for Arabic readability assessment. BAREC consists of 69,441 sentences spanning 1+ million words, carefully curated to cover…

计算与语言 · 计算机科学 2025-06-17 Khalid N. Elmadani , Nizar Habash , Hanada Taha-Thomure

We introduce the Tarab Corpus, a large-scale cultural and linguistic resource that brings together Arabic song lyrics and poetry within a unified analytical framework. The corpus comprises 2.56 million verses and more than 13.5 million…

计算与语言 · 计算机科学 2026-03-18 Mo El-Haj

Arabic is recognised as the 4th most used language of the Internet. Arabic has three main varieties: (1) classical Arabic (CA), (2) Modern Standard Arabic (MSA), (3) Arabic Dialect (AD). MSA and AD could be written either in Arabic or in…

计算与语言 · 计算机科学 2019-03-08 Imane Guellil , Houda Saâdane , Faical Azouaou , Billel Gueni , Damien Nouvel

This paper describes a web-based corpus of global language use with a focus on how this corpus can be used for data-driven language mapping. First, the corpus provides a representation of where national varieties of major languages are used…

计算与语言 · 计算机科学 2020-04-03 Jonathan Dunn

There are many known Arabic lexicons organized on different ways, each of them has a different number of Arabic words according to its organization way. This paper has used mathematical relations to count a number of Arabic words, which…

计算与语言 · 计算机科学 2013-11-26 Nidhal El-Abbadi , Ahmed Nidhal Khdhair , Adel Al-Nasrawi

In this paper we present the final result of a project on Tunisian Arabic encoded in Arabizi, the Latin-based writing system for digital conversations. The project led to the creation of two integrated and independent resources: a corpus…

计算与语言 · 计算机科学 2022-07-12 Elisa Gugliotta , Marco Dinarelli

We present the SAMER Corpus, the first manually annotated Arabic parallel corpus for text simplification targeting school-aged learners. Our corpus comprises texts of 159K words selected from 15 publicly available Arabic fiction novels most…

计算与语言 · 计算机科学 2024-04-30 Bashar Alhafni , Reem Hazim , Juan Piñeros Liberato , Muhamed Al Khalil , Nizar Habash

Text summarization has been intensively studied in many languages, and some languages have reached advanced stages. Yet, Arabic Text Summarization (ATS) is still in its developing stages. Existing ATS datasets are either small or lack…

计算与语言 · 计算机科学 2022-10-26 Abdulaziz Alhamadani , Xuchao Zhang , Jianfeng He , Chang-Tien Lu

Diacritization of Arabic text is both an interesting and a challenging problem at the same time with various applications ranging from speech synthesis to helping students learning the Arabic language. Like many other tasks or problems in…

计算与语言 · 计算机科学 2019-05-07 Ali Fadel , Ibraheem Tuffaha , Bara' Al-Jawarneh , Mahmoud Al-Ayyoub

Arabic language is one of the most popular languages in the world. Hundreds of millions of people in many countries around the world speak Arabic as their native speaking. However, due to complexity of Arabic language, recognition of…

计算机视觉与模式识别 · 计算机科学 2017-02-07 Ashraf A. Shahin

We introduce the largest transcribed Arabic speech corpus, QASR, collected from the broadcast domain. This multi-dialect speech dataset contains 2,000 hours of speech sampled at 16kHz crawled from Aljazeera news channel. The dataset is…

计算与语言 · 计算机科学 2021-06-25 Hamdy Mubarak , Amir Hussein , Shammur Absar Chowdhury , Ahmed Ali

We present the QuranMorph corpus, a morphologically annotated corpus for the Quran (77,429 tokens). Each token in the QuranMorph was manually lemmatized and tagged with its part-of-speech by three expert linguists. The lemmatization process…

计算与语言 · 计算机科学 2025-06-24 Diyam Akra , Tymaa Hammouda , Mustafa Jarrar

Automatic readability assessment is relevant to building NLP applications for education, content analysis, and accessibility. However, Arabic readability assessment is a challenging task due to Arabic's morphological richness and limited…

计算与语言 · 计算机科学 2024-07-04 Juan Piñeros Liberato , Bashar Alhafni , Muhamed Al Khalil , Nizar Habash

The ambition of a character recognition system is to transform a text document typed on paper into a digital format that can be manipulated by word processor software Unlike other languages, Arabic has unique features, while other language…

计算与语言 · 计算机科学 2010-06-15 A. A Zaidan , B. B Zaidan , Hamid. A. Jalab , Hamdan. O. Alanazi , Rami Alnaqeib

Language models (LMs) have introduced a major paradigm shift in Natural Language Processing (NLP) modeling where large pre-trained LMs became integral to most of the NLP tasks. The LMs are intelligent enough to find useful and relevant…

计算与语言 · 计算机科学 2023-05-09 Abbas Raza Ali , Muhammad Ajmal Siddiqui , Rema Algunaibet , Hasan Raza Ali

We present a formal Arabic wordnet built on the basis of a carefully designed ontology hereby referred to as the Arabic Ontology. The ontology provides a formal representation of the concepts that the Arabic terms convey, and its content…

计算与语言 · 计算机科学 2022-05-20 Mustafa Jarrar

In recent years, Large Language Models have revolutionized the field of natural language processing, showcasing an impressive rise predominantly in English-centric domains. These advancements have set a global benchmark, inspiring…

计算与语言 · 计算机科学 2024-05-06 Manel Aloui , Hasna Chouikhi , Ghaith Chaabane , Haithem Kchaou , Chehir Dhaouadi
‹ 上一页 1 2 3 10 下一页 ›