中文
相关论文

相关论文: ADAB: Arabic Dataset for Automated Politeness Benc…

200 篇论文

Natural Language Processing (NLP) is today a very active field of research and innovation. Many applications need however big sets of data for supervised learning, suitably labelled for the training purpose. This includes applications for…

计算与语言 · 计算机科学 2021-02-23 ElMehdi Boujou , Hamza Chataoui , Abdellah El Mekki , Saad Benjelloun , Ikram Chairi , Ismail Berrada

Despite its significance, Arabic, a linguistically rich and morphologically complex language, faces the challenge of being under-resourced. The scarcity of large annotated datasets hampers the development of accurate tools for subjectivity…

计算与语言 · 计算机科学 2026-03-02 Slimane Bellaouar , Attia Nehar , Soumia Souffi , Mounia Bouameur

The NLP pipeline has evolved dramatically in the last few years. The first step in the pipeline is to find suitable annotated datasets to evaluate the tasks we are trying to solve. Unfortunately, most of the published datasets lack metadata…

计算与语言 · 计算机科学 2021-10-14 Zaid Alyafeai , Maraim Masoud , Mustafa Ghaleb , Maged S. Al-shaibani

When building NLP models, there is a tendency to aim for broader coverage, often overlooking cultural and (socio)linguistic nuance. In this position paper, we make the case for care and attention to such nuances, particularly in dataset…

计算与语言 · 计算机科学 2022-03-21 A. Stevie Bergman , Mona T. Diab

In this paper, we present the annotation pipeline and the guidelines we wrote as part of an effort to create a large manually annotated Arabic author profiling dataset from various social media sources covering 16 Arabic countries and 11…

计算与语言 · 计算机科学 2018-08-24 Wajdi Zaghouani , Anis Charfi

We introduce ALARB, a dataset and suite of tasks designed to evaluate the reasoning capabilities of large language models (LLMs) within the Arabic legal domain. While existing Arabic benchmarks cover some knowledge-intensive tasks such as…

In this paper, we present Arap-Tweet, which is a large-scale and multi-dialectal corpus of Tweets from 11 regions and 16 countries in the Arab world representing the major Arabic dialectal varieties. To build this corpus, we collected data…

计算与语言 · 计算机科学 2018-08-24 Wajdi Zaghouani , Anis Charfi

We study politeness phenomena in nine typologically diverse languages. Politeness is an important facet of communication and is sometimes argued to be cultural-specific, yet existing computational linguistic study is limited to English. We…

计算与语言 · 计算机科学 2022-11-30 Anirudh Srinivasan , Eunsol Choi

Arabic, with its rich diversity of dialects, remains significantly underrepresented in Large Language Models, particularly in dialectal variations. We address this gap by introducing seven synthetic datasets in dialects alongside Modern…

We present our effort to create a large Multi-Layered representational repository of Linguistic Code-Switched Arabic data. The process involves developing clear annotation standards and Guidelines, streamlining the annotation process, and…

计算与语言 · 计算机科学 2019-10-01 Mona Diab , Mahmoud Ghoneim , Abdelati Hawwari , Fahad AlGhamdi , Nada AlMarwani , Mohamed Al-Badrashiny

The detection of toxic language in the Arabic language has emerged as an active area of research in recent years, and reviewing the existing datasets employed for training the developed solutions has become a pressing need. This paper…

计算与语言 · 计算机科学 2024-01-31 Imene Bensalem , Paolo Rosso , Hanane Zitouni

On annotating multi-dialect Arabic datasets, it is common to randomly assign the samples across a pool of native Arabic speakers. Recent analyses recommended routing dialectal samples to native speakers of their respective dialects to build…

计算与语言 · 计算机科学 2024-06-10 Amr Keleg , Walid Magdy , Sharon Goldwater

Since their inception, transformer-based language models have led to impressive performance gains across multiple natural language processing tasks. For Arabic, the current state-of-the-art results on most datasets are achieved by the…

计算与语言 · 计算机科学 2021-03-11 Amey Hengle , Atharva Kshirsagar , Shaily Desai , Manisha Marathe

With the rise of digital communication, memes have become a significant medium for cultural and political expression that is often used to mislead audiences. Identification of such misleading and persuasive multimodal content has become…

计算与语言 · 计算机科学 2024-10-08 Firoj Alam , Abul Hasnat , Fatema Ahmed , Md Arid Hasan , Maram Hasanain

The hospitality industry in the Arab world increasingly relies on customer feedback to shape services, driving the need for advanced Arabic sentiment analysis tools. To address this challenge, the Sentiment Analysis on Arabic Dialects in…

计算与语言 · 计算机科学 2025-11-18 Maram Alharbi , Salmane Chafik , Saad Ezzini , Ruslan Mitkov , Tharindu Ranasinghe , Hansi Hettiarachchi

The Arabic language is among the most popular languages in the world with a huge variety of dialects spoken in 22 countries. In this study, we address the problem of classifying 18 Arabic dialects of the QADI dataset of Arabic tweets. RNN…

计算与语言 · 计算机科学 2025-07-01 Omar A. Essameldin , Ali O. Elbeih , Wael H. Gomaa , Wael F. Elsersy

We present QADI, an automatically collected dataset of tweets belonging to a wide range of country-level Arabic dialects -covering 18 different countries in the Middle East and North Africa region. Our method for building this dataset…

计算与语言 · 计算机科学 2020-05-18 Ahmed Abdelali , Hamdy Mubarak , Younes Samih , Sabit Hassan , Kareem Darwish

There is a significant gap in evaluating cultural reasoning in LLMs using conversational datasets that capture culturally rich and dialectal contexts. Most Arabic benchmarks focus on short text snippets in Modern Standard Arabic (MSA),…

While recent Arabic NLP benchmarks focus on scale, they often rely on synthetic or translated data which may benefit from deeper linguistic verification. We introduce ALPS (Arabic Linguistic & Pragmatic Suite), a native, expert-curated…

计算与语言 · 计算机科学 2026-02-20 Hussein S. Al-Olimat , Ahmad Alshareef
‹ 上一页 1 2 3 10 下一页 ›