中文
相关论文

相关论文: EveTAR: Building a Large-Scale Multi-Task Test Col…

200 篇论文

Building a benchmark dataset for hate speech detection presents various challenges. Firstly, because hate speech is relatively rare, random sampling of tweets to annotate is very inefficient in finding hate speech. To address this, prior…

计算与语言 · 计算机科学 2021-11-11 Md Mustafizur Rahman , Dinesh Balakrishnan , Dhiraj Murthy , Mucahid Kutlu , Matthew Lease

Detecting offensive language on Twitter has many applications ranging from detecting/predicting bullying to measuring polarization. In this paper, we focus on building a large Arabic offensive tweet dataset. We introduce a method for…

计算与语言 · 计算机科学 2021-03-11 Hamdy Mubarak , Ammar Rashed , Kareem Darwish , Younes Samih , Ahmed Abdelali

Events detected from social media streams often include early signs of accidents, crimes or disasters. Therefore, they can be used by related parties for timely and efficient response. Although significant progress has been made on event…

社会与信息网络 · 计算机科学 2020-02-12 Yi Han , Shanika Karunasekera , Christopher Leckie

We aim at solving the problem of predicting people's ideology, or political tendency. We estimate it by using Twitter data, and formalize it as a classification problem. Ideology-detection has long been a challenging yet important problem.…

机器学习 · 计算机科学 2020-06-19 Zhiping Xiao , Weiping Song , Haoyan Xu , Zhicheng Ren , Yizhou Sun

The experimental landscape in natural language processing for social media is too fragmented. Each year, new shared tasks and datasets are proposed, ranging from classics like sentiment analysis to irony detection or emoji prediction.…

计算与语言 · 计算机科学 2020-10-27 Francesco Barbieri , Jose Camacho-Collados , Leonardo Neves , Luis Espinosa-Anke

Social media has become an important information source for crisis management and provides quick access to ongoing developments and critical information. However, classification models suffer from event-related biases and highly imbalanced…

计算与语言 · 计算机科学 2022-11-22 Philipp Seeberger , Korbinian Riedhammer

The complexities of Arabic language in morphology, orthography and dialects makes sentiment analysis for Arabic more challenging. Also, text feature extraction from short messages like tweets, in order to gauge the sentiment, makes this…

计算与语言 · 计算机科学 2018-10-17 Abdulaziz M. Alayba , Vasile Palade , Matthew England , Rahat Iqbal

During the onset of a disaster event, filtering relevant information from the social web data is challenging due to its sparse availability and practical limitations in labeling datasets of an ongoing crisis. In this paper, we hypothesize…

计算与语言 · 计算机科学 2020-10-22 Jitin Krishnan , Hemant Purohit , Huzefa Rangwala

Twitter stream has become a large source of information for many people, but the magnitude of tweets and the noisy nature of its content have made harvesting the knowledge from Twitter a challenging task for researchers for a long time.…

计算与语言 · 计算机科学 2018-06-21 Øystein Repp , Heri Ramampiaro

Tweets pertaining to a single event, such as a national election, can number in the hundreds of millions. Automatically analyzing them is beneficial in many downstream natural language applications such as question answering and…

计算与语言 · 计算机科学 2013-11-06 Saif M. Mohammad , Svetlana Kiritchenko , Joel Martin

Social media plays a significant role in disaster management by providing valuable data about affected people, donations and help requests. Recent studies highlight the need to filter information on social media into fine-grained content…

计算与语言 · 计算机科学 2021-05-20 Hamada M. Zahera , Rricha Jalota , Mohamed A. Sherif , Axel N. Ngomo

Twitter, a microblogging service, is todays most popular platform for communication in the form of short text messages, called Tweets. Users use Twitter to publish their content either for expressing concerns on information news or views on…

社会与信息网络 · 计算机科学 2017-11-29 Dhanasekar Sundararaman , Priya Arora , Vishwanath Seshagiri

To create a new IR test collection at low cost, it is valuable to carefully select which documents merit human relevance judgments. Shared task campaigns such as NIST TREC pool document rankings from many participating systems (and often…

信息检索 · 计算机科学 2020-08-06 Md Mustafizur Rahman , Mucahid Kutlu , Tamer Elsayed , Matthew Lease

Social media classification tasks (e.g., tweet sentiment analysis, tweet stance detection) are challenging because social media posts are typically short, informal, and ambiguous. Thus, training on tweets is challenging and demands…

计算与语言 · 计算机科学 2023-02-21 Shizhe Diao , Sedrick Scott Keh , Liangming Pan , Zhiliang Tian , Yan Song , Tong Zhang

We present QADI, an automatically collected dataset of tweets belonging to a wide range of country-level Arabic dialects -covering 18 different countries in the Middle East and North Africa region. Our method for building this dataset…

计算与语言 · 计算机科学 2020-05-18 Ahmed Abdelali , Hamdy Mubarak , Younes Samih , Sabit Hassan , Kareem Darwish

In this paper, we provide an overview of the WNUT-2020 shared task on the identification of informative COVID-19 English Tweets. We describe how we construct a corpus of 10K Tweets and organize the development and evaluation phases for this…

计算与语言 · 计算机科学 2020-10-19 Dat Quoc Nguyen , Thanh Vu , Afshin Rahimi , Mai Hoang Dao , Linh The Nguyen , Long Doan

In this paper, we tackle the Arabic Fine-Grained Hate Speech Detection shared task and demonstrate significant improvements over reported baselines for its three subtasks. The tasks are to predict if a tweet contains (1) Offensive language;…

计算与语言 · 计算机科学 2022-05-18 Badr AlKhamissi , Mona Diab

Veracity of data posted on the microblog platforms has in recent years been a subject of intensive study by professionals specializing in various fields of informatics as well as sociology, particularly in the light of increasing importance…

社会与信息网络 · 计算机科学 2021-01-20 Majed Alrubaian , Muhammad Al-Qurishi , Sherif Omar , Mohamed A. Mostafa

We describe the fourth edition of the CheckThat! Lab, part of the 2021 Conference and Labs of the Evaluation Forum (CLEF). The lab evaluates technology supporting tasks related to factuality, and covers Arabic, Bulgarian, English, Spanish,…

This paper provides an overview of the Arabic Sentiment Analysis Challenge organized by King Abdullah University of Science and Technology (KAUST). The task in this challenge is to develop machine learning models to classify a given tweet…