中文
相关论文

相关论文: The Evolution of Darija Open Dataset: Introducing …

200 篇论文

Arabic is a linguistically and culturally rich language with a vast vocabulary that spans scientific, religious, and literary domains. Yet, large-scale lexical datasets linking Arabic words to precise definitions remain limited. We present…

计算与语言 · 计算机科学 2026-01-30 Serry Sibaee , Yasser Alhabashi , Nadia Sibai , Yara Farouk , Adel Ammar , Sawsan AlHalawani , Wadii Boulila

This survey offers a comprehensive overview of Large Language Models (LLMs) designed for Arabic language and its dialects. It covers key architectures, including encoder-only, decoder-only, and encoder-decoder models, along with the…

计算与语言 · 计算机科学 2026-05-20 Malak Mashaabi , Shahad Al-Khalifa , Hend Al-Khalifa

In recent years, Large Language Models have revolutionized the field of natural language processing, showcasing an impressive rise predominantly in English-centric domains. These advancements have set a global benchmark, inspiring…

计算与语言 · 计算机科学 2024-05-06 Manel Aloui , Hasna Chouikhi , Ghaith Chaabane , Haithem Kchaou , Chehir Dhaouadi

In recent years, the enhanced capabilities of ASR models and the emergence of multi-dialect datasets have increasingly pushed Arabic ASR model development toward an all-dialect-in-one direction. This trend highlights the need for…

计算与语言 · 计算机科学 2024-12-19 Yingzhi Wang , Anas Alhmoud , Muhammad Alqurishi

Arabic is a complex language with many varieties and dialects spoken by over 450 millions all around the world. Due to the linguistic diversity and variations, it is challenging to build a robust and generalized ASR system for Arabic. In…

计算与语言 · 计算机科学 2023-10-30 Abdul Waheed , Bashar Talafha , Peter Sullivan , AbdelRahim Elmadany , Muhammad Abdul-Mageed

The rapid digitalization of customer service has intensified the demand for conversational agents capable of providing accurate and natural interactions. In the Algerian context, this is complicated by the linguistic complexity of Darja, a…

计算与语言 · 计算机科学 2026-02-03 El Batoul Bechiri , Dihia Lanasri

The size of large language models (LLMs) has scaled dramatically in recent years and their computational and data requirements have surged correspondingly. State-of-the-art language models, even at relatively smaller sizes, typically…

This paper addresses the critical need for democratizing large language models (LLM) in the Arab world, a region that has seen slower progress in developing models comparable to state-of-the-art offerings like GPT-4 or ChatGPT 3.5, due to a…

This paper presents LOLA, a massively multilingual large language model trained on more than 160 languages using a sparse Mixture-of-Experts Transformer architecture. Our architectural and implementation choices address the challenge of…

Arabic, with its rich diversity of dialects, remains significantly underrepresented in Large Language Models, particularly in dialectal variations. We address this gap by introducing seven synthetic datasets in dialects alongside Modern…

Dialectal Arabic is the primary spoken language used by native Arabic speakers in daily communication. The rise of social media platforms has notably expanded its use as a written language. However, Arabic dialects do not have standard…

计算与语言 · 计算机科学 2025-06-02 Bashar Alhafni , Sarah Al-Towaity , Ziyad Fawzy , Fatema Nassar , Fadhl Eryani , Houda Bouamor , Nizar Habash

With the advent of informal electronic communications such as social media, colloquial languages that were historically unwritten are being written for the first time in heavily code-switched environments. We present a method for inducing…

计算与语言 · 计算机科学 2017-08-22 Michael Bloodgood , Benjamin Strauss

Due to their crucial role in all NLP, several benchmarks have been proposed to evaluate pretrained language models. In spite of these efforts, no public benchmark of diverse nature currently exists for evaluation of Arabic. This makes it…

计算与语言 · 计算机科学 2023-05-31 AbdelRahim Elmadany , El Moatez Billah Nagoudi , Muhammad Abdul-Mageed

Multilingual task-oriented dialogue (ToD) facilitates access to services and information for many (communities of) speakers. Nevertheless, the potential of this technology is not fully realised, as current datasets for multilingual ToD -…

计算与语言 · 计算机科学 2023-05-24 Olga Majewska , Evgeniia Razumovskaia , Edoardo Maria Ponti , Ivan Vulić , Anna Korhonen

The NLP pipeline has evolved dramatically in the last few years. The first step in the pipeline is to find suitable annotated datasets to evaluate the tasks we are trying to solve. Unfortunately, most of the published datasets lack metadata…

计算与语言 · 计算机科学 2021-10-14 Zaid Alyafeai , Maraim Masoud , Mustafa Ghaleb , Maged S. Al-shaibani

Translation between natural language and source code can help software development by enabling developers to comprehend, ideate, search, and write computer programs in natural language. Despite growing interest from the industry and the…

The growing importance of culturally-aware natural language processing systems has led to an increasing demand for resources that capture sociopragmatic phenomena across diverse languages. Nevertheless, Arabic-language resources for…

Information about pretraining corpora used to train the current best-performing language models is seldom discussed: commercial models rarely detail their data, and even open models are often released without accompanying training data or…

The hospitality industry in the Arab world increasingly relies on customer feedback to shape services, driving the need for advanced Arabic sentiment analysis tools. To address this challenge, the Sentiment Analysis on Arabic Dialects in…

计算与语言 · 计算机科学 2025-11-18 Maram Alharbi , Salmane Chafik , Saad Ezzini , Ruslan Mitkov , Tharindu Ranasinghe , Hansi Hettiarachchi

Existing large language models (LLMs) that mainly focus on Standard American English (SAE) often lead to significantly worse performance when being applied to other English dialects. While existing mitigations tackle discrepancies for…

计算与语言 · 计算机科学 2023-12-07 Yanchen Liu , William Held , Diyi Yang