中文
相关论文

相关论文: 1.5 billion words Arabic Corpus

200 篇论文

International library standards require cataloguers to tediously input Romanization of their catalogue records for the benefit of library users without specific language expertise. In this paper, we present the first reported results on the…

计算与语言 · 计算机科学 2021-03-15 Eryani Fadhl , Habash Nizar

Cross-Lingual SynthDocs is a large-scale synthetic corpus designed to address the scarcity of Arabic resources for Optical Character Recognition (OCR) and Document Understanding (DU). The dataset comprises over 2.5 million of samples,…

This work consists of creating a system of the Computer Assisted Language Learning (CALL) based on a system of Automatic Speech Recognition (ASR) for the Arabic language using the tool CMU Sphinx3 [1], based on the approach of HMM. To this…

计算与语言 · 计算机科学 2012-05-16 Naim Terbeh , Mounir Zrigui

The complexities of Arabic language in morphology, orthography and dialects makes sentiment analysis for Arabic more challenging. Also, text feature extraction from short messages like tweets, in order to gauge the sentiment, makes this…

计算与语言 · 计算机科学 2018-10-17 Abdulaziz M. Alayba , Vasile Palade , Matthew England , Rahat Iqbal

The availability of corpora is a major factor in building natural language processing applications. However, the costs of acquiring corpora can prevent some researchers from going further in their endeavours. The ease of access to freely…

计算与语言 · 计算机科学 2017-02-28 Wajdi Zaghouani

We present the IndicNLP corpus, a large-scale, general-domain corpus containing 2.7 billion words for 10 Indian languages from two language families. We share pre-trained word embeddings trained on these corpora. We create news article…

The rapid growth of the internet has increased the number of online texts. This led to the rapid growth of the number of online texts in the Arabic language. The enormous amount of text must be organized into classes to make the analysis…

信息检索 · 计算机科学 2022-11-08 Sumaia Mohammed AL-Ghuribi , Shahrul Azman Mohd Noah

In this work, we address the problem of spelling correction in the Arabic language utilizing the new corpus provided by QALB (Qatar Arabic Language Bank) project which is an annotated corpus of sentences with errors and their corrections.…

机器学习 · 计算机科学 2014-10-01 Youssef Hassan , Mohamed Aly , Amir Atiya

Gender bias in natural language processing (NLP) applications, particularly machine translation, has been receiving increasing attention. Much of the research on this issue has focused on mitigating gender bias in English NLP models and…

计算与语言 · 计算机科学 2021-10-19 Bashar Alhafni , Nizar Habash , Houda Bouamor

With the development of electronic media and the heterogeneity of Arabic data on the Web, the idea of building a clean corpus for certain applications of natural language processing, including machine translation, information retrieval,…

计算与语言 · 计算机科学 2017-09-28 Wided Bakari , Patrice Bellot , Mahmoud Neji

Over the past years, interest in discourse analysis and discourse parsing has steadily grown, and many discourse-annotated corpora and, as a result, discourse parsers have been built. In this paper, we present a discourse-annotated corpus…

计算与语言 · 计算机科学 2021-06-29 Sara Shahmohammadi , Hadi Veisi , Ali Darzi

This paper presents a dataset for closest opposite questions in Arabic language. The dataset is the first of its kind for the Arabic language. It is beneficial for the assessment of systems on the aspect of antonymy detection. The structure…

计算与语言 · 计算机科学 2023-10-24 Sandra Rizkallah , Amir F. Atiya , Samir Shaheen

Pre-trained Language Models (PLMs) are integral to many modern natural language processing (NLP) systems. Although multilingual models cover a wide range of languages, they often grapple with challenges like high inference costs and a lack…

计算与语言 · 计算机科学 2024-07-19 Murtadha Ahmed , Saghir Alfasly , Bo Wen , Jamaal Qasem , Mohammed Ahmed , Yunfeng Liu

The study of natural language, especially Arabic, and mechanisms for the implementation of automatic processing is a fascinating field of study, with various potential applications. The importance of tools for natural language processing is…

计算与语言 · 计算机科学 2013-06-05 Riadh Bouslimi , Houda Amraoui

Kurdish is a less-resourced language consisting of different dialects written in various scripts. Approximately 30 million people in different countries speak the language. The lack of corpora is one of the main obstacles in Kurdish…

计算与语言 · 计算机科学 2019-09-26 Roshna Omer Abdulrahman , Hossein Hassani , Sina Ahmadi

In this paper, we present the annotation pipeline and the guidelines we wrote as part of an effort to create a large manually annotated Arabic author profiling dataset from various social media sources covering 16 Arabic countries and 11…

计算与语言 · 计算机科学 2018-08-24 Wajdi Zaghouani , Anis Charfi

We present Qabas, a novel open-source Arabic lexicon designed for NLP applications. The novelty of Qabas lies in its synthesis of 110 lexicons. Specifically, Qabas lexical entries (lemmas) are assembled by linking lemmas from 110 lexicons.…

计算与语言 · 计算机科学 2024-06-12 Mustafa Jarrar , Tymaa Hammouda

Informal language is a style of spoken or written language frequently used in casual conversations, social media, weblogs, emails and text messages. In informal writing, the language faces some lexical and/or syntactic changes varying among…

计算与语言 · 计算机科学 2023-08-11 Vahide Tajalli , Fateme Kalantari , Mehrnoush Shamsfard

In this paper, a supervised learning technique for extracting keyphrases of Arabic documents is presented. The extractor is supplied with linguistic knowledge to enhance its efficiency instead of relying only on statistical information such…

计算与语言 · 计算机科学 2012-03-22 Tarek El-shishtawy , Abdulwahab Al-sammak

This study presents EgyBERT, an Arabic language model pretrained on 10.4 GB of Egyptian dialectal texts. We evaluated EgyBERT's performance by comparing it with five other multidialect Arabic language models across 10 evaluation datasets.…

计算与语言 · 计算机科学 2024-08-08 Faisal Qarah