English
Related papers

Related papers: Empirical Evaluation of Public HateSpeech Datasets

200 papers

While social media empowers freedom of expression and individual voices, it also enables anti-social behavior, online harassment, cyberbullying, and hate speech. In this paper, we deepen our understanding of online hate speech by focusing…

Computation and Language · Computer Science 2018-04-13 Mai ElSherief , Vivek Kulkarni , Dana Nguyen , William Yang Wang , Elizabeth Belding

Our study addresses a significant gap in online hate speech detection research by focusing on homophobia, an area often neglected in sentiment analysis research. Utilising advanced sentiment analysis models, particularly BERT, and…

Computation and Language · Computer Science 2024-05-16 Josh McGiff , Nikola S. Nikolov

In this paper we examine methods to detect hate speech in social media, while distinguishing this from general profanity. We aim to establish lexical baselines for this task by applying supervised classification methods using a recently…

Computation and Language · Computer Science 2017-12-29 Shervin Malmasi , Marcos Zampieri

Hateful memes pose a unique challenge for current machine learning systems because their message is derived from both text- and visual-modalities. To this effect, Facebook released the Hateful Memes Challenge, a dataset of memes with…

Computer Vision and Pattern Recognition · Computer Science 2021-07-12 Hannah Rose Kirk , Yennie Jun , Paulius Rauba , Gal Wachtel , Ruining Li , Xingjian Bai , Noah Broestl , Martin Doff-Sotta , Aleksandar Shtedritski , Yuki M. Asano

Online social media is rife with offensive and hateful comments, prompting the need for their automatic detection given the sheer amount of posts created every second. Creating high-quality human-labelled datasets for this task is difficult…

Computation and Language · Computer Science 2023-08-01 João A. Leite , Carolina Scarton , Diego F. Silva

Hate speech detection models are only as good as the data they are trained on. Datasets sourced from social media suffer from systematic gaps and biases, leading to unreliable models with simplistic decision boundaries. Adversarial…

Computation and Language · Computer Science 2024-03-29 Janis Goldzycher , Paul Röttger , Gerold Schneider

Counterspeech, i.e., responses to counteract potential harms of hateful speech, has become an increasingly popular solution to address online hate speech without censorship. However, properly countering hateful language requires countering…

Computation and Language · Computer Science 2023-11-02 Jimin Mun , Emily Allaway , Akhila Yerukola , Laura Vianna , Sarah-Jane Leslie , Maarten Sap

In recent years, the rising use of social media has propelled automated cyberbullying detection into a prominent research domain. However, challenges persist due to the absence of a standardized definition and universally accepted datasets.…

Social and Information Networks · Computer Science 2024-02-19 Naveed Ejaz , Fakhra Kashif , Salimur Choudhury

Toxic speech, also known as hate speech, is regarded as one of the crucial issues plaguing online social media today. Most recent work on toxic speech detection is constrained to the modality of text and written conversations with very…

Computation and Language · Computer Science 2022-04-05 Sreyan Ghosh , Samden Lepcha , S Sakshi , Rajiv Ratn Shah , S. Umesh

Community models for malicious content detection, which take into account the context from a social graph alongside the content itself, have shown remarkable performance on benchmark datasets. Yet, misinformation and hate speech continue to…

Machine Learning · Computer Science 2024-09-30 Ivo Verhoeven , Pushkar Mishra , Rahel Beloch , Helen Yannakoudakis , Ekaterina Shutova

Hateful comments are prevalent on social media platforms. Although tools for automatically detecting, flagging, and blocking such false, offensive, and harmful content online have lately matured, such reactive and brute force methods alone…

Computation and Language · Computer Science 2024-01-17 Sougata Saha , Rohini Srihari

The growing use of social media has led to the development of several Machine Learning (ML) and Natural Language Processing(NLP) tools to process the unprecedented amount of social media content to make actionable decisions. However, these…

Computation and Language · Computer Science 2021-10-28 Izzat Alsmadi , Kashif Ahmad , Mahmoud Nazzal , Firoj Alam , Ala Al-Fuqaha , Abdallah Khreishah , Abdulelah Algosaibi

Automatic hate speech detection is hampered by the scarcity of labeled datasetd, leading to poor generalization. We employ pretrained language models (LMs) to alleviate this data bottleneck. We utilize the GPT LM for generating large…

Computation and Language · Computer Science 2021-09-03 Tomer Wullach , Amir Adler , Einat Minkov

While social media has become an empowering agent to individual voices and freedom of expression, it also facilitates anti-social behaviors including online harassment, cyberbullying, and hate speech. In this paper, we present the first…

Social and Information Networks · Computer Science 2018-04-13 Mai ElSherief , Shirin Nilizadeh , Dana Nguyen , Giovanni Vigna , Elizabeth Belding

Online social networks have become a fundamental component of our everyday life. Unfortunately, these platforms are also a stage for hate speech. Popular social networks have regularized rules against hate speech. Consequently, social…

Computation and Language · Computer Science 2023-04-04 Nadav Schneider , Shimon Shouei , Saleem Ghantous , Elad Feldman

Hate speech detection is a challenging problem with most of the datasets available in only one language: English. In this paper, we conduct a large scale analysis of multilingual hate speech in 9 languages from 16 different sources. We…

Social and Information Networks · Computer Science 2020-12-10 Sai Saketh Aluru , Binny Mathew , Punyajoy Saha , Animesh Mukherjee

We examined four case studies in the context of hate speech on Twitter in Italian from 2019 to 2020, aiming at comparing the classification of the 3,600 tweets made by expert pedagogists with the automatic classification made by machine…

Social and Information Networks · Computer Science 2024-02-14 Erica Forzinetti , Marco L. Della Vedova , Stefano Pasta , Milena Santerini

This study uses the cosine similarity ratio, embedding regression, and manual re-annotation to diagnose hate speech classification. We begin by computing cosine similarity ratio on a dataset "Measuring Hate Speech" that contains 135,556…

Computation and Language · Computer Science 2024-11-27 Xilin Yang

Current research in machine learning and artificial intelligence is largely centered on modeling and performance evaluation, less so on data collection. However, recent research demonstrated that limitations and biases in data may…

Artificial Intelligence · Computer Science 2025-02-18 Eleonora Mancini , Ana Tanevska , Andrea Galassi , Alessio Galatolo , Federico Ruggeri , Paolo Torroni

Online harassment has been a problem to a greater or lesser extent since the early days of the internet. Previous work has applied anti-spam techniques like machine-learning based text classification (Reynolds, 2011) to detecting harassing…

Computation and Language · Computer Science 2016-09-12 Alexei Bastidas , Edward Dixon , Chris Loo , John Ryan