中文
相关论文

相关论文: Algerian Dialect

200 篇论文

With the recent proliferation of open textual data on social media platforms, Emotion Detection (ED) from Text has received more attention over the past years. It has many applications, especially for businesses and online service…

计算与语言 · 计算机科学 2022-07-26 Hossein Mirzaee , Javad Peymanfard , Hamid Habibzadeh Moshtaghin , Hossein Zeinali

This article presents morphologically-annotated Yemeni, Sudanese, Iraqi, and Libyan Arabic dialects Lisan corpora. Lisan features around 1.2 million tokens. We collected the content of the corpora from several social media platforms. The…

计算与语言 · 计算机科学 2022-12-20 Mustafa Jarrar , Fadi A Zaraket , Tymaa Hammouda , Daanish Masood Alavi , Martin Waahlisch

Abusive language is a massive problem in online social platforms. Existing abusive language detection techniques are particularly ill-suited to comments containing heterogeneous abusive language patterns, i.e., both abusive and non-abusive…

计算与语言 · 计算机科学 2021-05-25 Hongyu Gong , Alberto Valido , Katherine M. Ingram , Giulia Fanti , Suma Bhat , Dorothy L. Espelage

Sentiment Analysis, a popular subtask of Natural Language Processing, employs computational methods to extract sentiment, opinions, and other subjective aspects from linguistic data. Given its crucial role in understanding human sentiment,…

计算与语言 · 计算机科学 2025-02-07 Zhiqiang Shi , Ruchit Agrawal

The NLP pipeline has evolved dramatically in the last few years. The first step in the pipeline is to find suitable annotated datasets to evaluate the tasks we are trying to solve. Unfortunately, most of the published datasets lack metadata…

计算与语言 · 计算机科学 2021-10-14 Zaid Alyafeai , Maraim Masoud , Mustafa Ghaleb , Maged S. Al-shaibani

Social media sites such as YouTube and Facebook have become an integral part of everyone's life and in the last few years, hate speech in the social media comment section has increased rapidly. Detection of hate speech on social media…

计算与语言 · 计算机科学 2020-12-18 Nauros Romim , Mosahed Ahmed , Hriteshwar Talukder , Md Saiful Islam

Arabic dialects form a diverse continuum, yet NLP models often treat them as discrete categories. Recent work addresses this issue by modeling dialectness as a continuous variable, notably through the Arabic Level of Dialectness (ALDi).…

计算与语言 · 计算机科学 2025-08-26 Sanad Shaban , Nizar Habash

Sentiment analysis for the Bengali language has attracted increasing research interest in recent years. However, progress remains constrained by the scarcity of large-scale and diverse annotated datasets. Although several Bengali sentiment…

计算与语言 · 计算机科学 2026-01-29 Akif Islam , Sujan Kumar Roy , Md. Ekramul Hamid

In this paper, we introduce the Akan Conversation Emotion (ACE) dataset, the first multimodal emotion dialogue dataset for an African language, addressing the significant lack of resources for low-resource languages in emotion recognition…

计算与语言 · 计算机科学 2025-06-04 David Sasu , Zehui Wu , Ziwei Gong , Run Chen , Pengyuan Shi , Lin Ai , Julia Hirschberg , Natalie Schluter

Natural language processing (NLP), particularly sentiment analysis, plays a vital role in areas like marketing, customer service, and social media monitoring by providing insights into user opinions and emotions. However, progress in Arabic…

计算与语言 · 计算机科学 2025-09-30 Dania Refai , Alaa Dalaq , Doaa Dalaq , Irfan Ahmad

Offensive language detection has been well studied in many languages, but it is lagging behind in low-resource languages, such as Hebrew. In this paper, we present a new offensive language corpus in Hebrew. A total of 15,881 tweets were…

计算与语言 · 计算机科学 2023-09-07 Nagham Hamad , Mustafa Jarrar , Mohammad Khalilia , Nadim Nashif

In the online world, Machine Translation (MT) systems are extensively used to translate User-Generated Text (UGT) such as reviews, tweets, and social media posts, where the main message is often the author's positive or negative attitude…

计算与语言 · 计算机科学 2023-06-09 Hadeel Saadany , Constantin Orasan , Emad Mohamed , Ashraf Tantawy

This research introduces a bilingual dataset comprising 23,456 entries for Arabic and 10,036 entries for English, annotated for emotions and hope speech, addressing the scarcity of multi-emotion (Emotion and hope) datasets. The dataset…

计算与语言 · 计算机科学 2025-05-22 Wajdi Zaghouani , Md. Rafiul Biswas

The problem of online offensive language limits the health and security of online users. It is essential to apply the latest state-of-the-art techniques in developing a system to detect online offensive language and to ensure social justice…

计算与语言 · 计算机科学 2022-03-08 Fatemah Husain , Ozlem Uzuner

The largest dataset of Arabic speech mispronunciation detections in Egyptian dialogues is introduced. The dataset is composed of annotated audio files representing the top 100 words that are most frequently used in the Arabic language,…

计算与语言 · 计算机科学 2021-11-03 Salah A. Aly , Abdelrahman Salah , Hesham M. Eraqi

Memes have become a prominent medium of political communication in the Arab world, reflecting how humor, imagery, and text interact to express ideological and cultural positions. Despite the centrality of memes to online political…

计算与语言 · 计算机科学 2026-05-21 Wajdi Zaghouani , Kais Attia , Md. Rafiul Biswas , Fadhl Eryani

The study of online discourse has become central to understanding societal polarization. While much research has focused on detecting overt toxicity, the subtle dynamics of social cohesion, meaning the interaction between divisive and…

计算与语言 · 计算机科学 2026-05-22 Aisha Ali Al-Athba , Wajdi Zaghouani

Investor sentiment shapes financial markets, yet modeling sentiment in Arabic financial contexts remains challenging due to linguistic complexity and limited resources. We present an Arabic NLP framework for large-scale financial sentiment…

计算与语言 · 计算机科学 2026-05-20 Mona H. Albaqawi , Eman M. Albalkhi , Joud A. Albaiti , Enrico Lopedoto

Linguistic uncertainty is a common feature of social media discourse, yet its relationship with user engagement remains underexplored, particularly in non-English contexts. Using a dataset of 16,695 Arabic-language tweets about Lebanon…

计算机与社会 · 计算机科学 2026-03-03 Mohamed Soufan

This study investigates logistic regression, linear support vector machine, multinomial Naive Bayes, and Bernoulli Naive Bayes for classifying Libyan dialect utterances gathered from Twitter. The dataset used is the QADI corpus, which…

计算与语言 · 计算机科学 2025-12-05 Mansour Essgaer , Khamis Massud , Rabia Al Mamlook , Najah Ghmaid