中文
相关论文

相关论文: 1.5 billion words Arabic Corpus

200 篇论文

In this study we analyzed a corpus of 8 million words academic literature from Computational lingustics' academic literature. the lexical bundles from this corpus are categorized based on structures and functions.

计算与语言 · 计算机科学 2016-03-11 Adel Rahimi

Enhancing existing models with new knowledge is a crucial aspect of AI development. This paper introduces a novel method for integrating a new language into a large language model (LLM). Our approach successfully incorporates a previously…

计算与语言 · 计算机科学 2025-08-22 Khalil Hennara , Sara Chrouf , Mohamed Motaism Hamed , Zeina Aldallal , Omar Hadid , Safwan AlModhayan

We present KS-PRET-5M, the largest publicly available pretraining dataset for the Kashmiri language, comprising 5,090,244 (5.09M) words, 27,692,959 (27.6M) characters, and a vocabulary of 295,433 (295.4K) unique word types. We assembled the…

计算与语言 · 计算机科学 2026-04-14 Haq Nawaz Malik , Nahfid Nissar

Natural language processing (NLP) utilizes text data augmentation to overcome sample size constraints. Increasing the sample size is a natural and widely used strategy for alleviating these challenges. In this study, we chose Arabic to…

计算与语言 · 计算机科学 2024-11-08 Ahlam Alrehili , Areej Alhothali

Automatic speech recognition (ASR) performs well for high-resource languages with abundant paired audio-transcript data, but its accuracy degrades sharply for most languages due to limited publicly available aligned data. To this end, we…

计算与语言 · 计算机科学 2026-05-12 Antonis Asonitis , Luca A. Lanzendörfer , Frédéric Berdoz , Roger Wattenhofer

We present Quran MD, a comprehensive multimodal dataset of the Quran that integrates textual, linguistic, and audio dimensions at the verse and word levels. For each verse (ayah), the dataset provides its original Arabic text, English…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Muhammad Umar Salman , Mohammad Areeb Qazi , Mohammed Talha Alam

Large language models (LLMs) have recently emerged as a powerful tool for a wide range of language generation tasks. Nevertheless, this progress has been slower in Arabic. In this work, we focus on the task of generating stories from LLMs.…

计算与语言 · 计算机科学 2024-07-11 Ahmed Oumar El-Shangiti , Fakhraddin Alwajih , Muhammad Abdul-Mageed

Displaying a document in Middle Eastern languages requires contextual analysis due to different presentational forms for each character of the alphabet. The words of the document will be formed by the joining of the correct positional…

计算与语言 · 计算机科学 2015-09-15 Kazem Taghva

Optical Character Recognition (OCR) is the process of extracting digitized text from images of scanned documents. While OCR systems have already matured in many languages, they still have shortcomings in cursive languages with overlapping…

计算机视觉与模式识别 · 计算机科学 2020-09-22 Hussein Osman , Karim Zaghw , Mostafa Hazem , Seifeldin Elsehely

With the rise of digital communication, memes have become a significant medium for cultural and political expression that is often used to mislead audiences. Identification of such misleading and persuasive multimodal content has become…

计算与语言 · 计算机科学 2024-10-08 Firoj Alam , Abul Hasnat , Fatema Ahmed , Md Arid Hasan , Maram Hasanain

Word embeddings are a core component of modern natural language processing systems, making the ability to thoroughly evaluate them a vital task. We describe DiaLex, a benchmark for intrinsic evaluation of dialectal Arabic word embedding.…

Text processing is one of the sub-branches of natural language processing. Recently, the use of machine learning and neural networks methods has been given greater consideration. For this reason, the representation of words has become very…

计算与语言 · 计算机科学 2017-12-20 Siamak Sarmady , Erfan Rahmani

The United States Code (Code) is a document containing over 22 million words that represents a large and important source of Federal statutory law. Scholars and policy advocates often discuss the direction and magnitude of changes in…

信息检索 · 计算机科学 2015-05-18 Michael J. Bommarito , Daniel Martin Katz

Large Language Models (LLMs) demonstrate remarkable fluency across high-resource languages yet consistently fail to generate coherent text in Kashmiri, a language spoken by approximately seven million people. This performance disparity…

计算与语言 · 计算机科学 2026-01-06 Haq Nawaz Malik

We present a collection of parallel corpora of 12 sign languages in video format, together with subtitles in the dominant spoken languages of the corresponding countries. The entire collection includes more than 1,300 hours in 4,381 video…

Question semantic similarity is a challenging and active research problem that is very useful in many NLP applications, such as detecting duplicate questions in community question answering platforms such as Quora. Arabic is considered to…

计算与语言 · 计算机科学 2019-09-23 Hesham Al-Bataineh , Wael Farhan , Ahmad Mustafa , Haitham Seelawi , Hussein T. Al-Natsheh

Despite growing interest in Quranic data research, existing Quran datasets remain limited in both scale and diversity. To address this gap, we present Tadabur, a large-scale Quran audio dataset. Tadabur comprises more than 1400+ hours of…

声音 · 计算机科学 2026-04-22 Faisal Alherran

Nowadays, it is no more needed to do an enormous effort to distribute a lot of forms to thousands of people and collect them, then convert this from into electronic format to track people opinion about some subjects. A lot of web sites can…

计算与语言 · 计算机科学 2020-01-29 Zitouni Abdelhafid , Hichem Rahab , Abdelhafid Zitouni , Mahieddine Djoudi

Code-mixing involves the seamless integration of linguistic elements from multiple languages within a single discourse, reflecting natural multilingual communication patterns. Despite its prominence in informal interactions such as social…

计算与语言 · 计算机科学 2025-06-17 Svetlana Churina , Akshat Gupta , Insyirah Mujtahid , Kokil Jaidka

This paper presents a novel dotless representation of Arabic text as an alternative to the standard Arabic text representation. We delve into its implications through comprehensive analysis across five diverse corpora and four different…

计算与语言 · 计算机科学 2023-12-27 Maged S. Al-Shaibani , Irfan Ahmad