English
Related papers

Related papers: SoftHateBench: Evaluating Moderation Models Agains…

200 papers

Hate speech is commonly defined as any communication that disparages a target group of people based on some characteristic such as race, colour, ethnicity, gender, sexual orientation, nationality, religion, or other characteristic. Due to…

Computation and Language · Computer Science 2018-09-13 Ona de Gibert , Naiara Perez , Aitor García-Pablos , Montse Cuadros

Hateful memes are an emerging method of spreading hate on the internet, relying on both images and text to convey a hateful message. We take an interpretable approach to hateful meme detection, using machine learning and simple heuristics…

Machine Learning · Computer Science 2021-08-24 Tanvi Deshpande , Nitya Mani

This paper investigates the use of machine learning models for the classification of unhealthy online conversations containing one or more forms of subtler abuse, such as hostility, sarcasm, and generalization. We leveraged a public dataset…

Computation and Language · Computer Science 2022-01-28 Shlok Gilda , Mirela Silva , Luiz Giovanini , Daniela Oliveira

Social media has a significant impact on people's lives. Hate speech on social media has emerged as one of society's most serious issues in recent years. Text and pictures are two forms of multimodal data that are distributed within…

Computation and Language · Computer Science 2024-09-18 Anusha Chhabra , Dinesh Kumar Vishwakarma

In the day and age of social media, users have become prone to online hate speech. Several attempts have been made to classify hate speech using machine learning but the state-of-the-art models are not robust enough for practical…

Computation and Language · Computer Science 2021-08-03 Tashvik Dhamija , Anjum , Rahul Katarya

Hate speech, offensive language, aggression, racism, sexism, and other abusive language are common phenomena in social media. There is a need for Artificial Intelligence(AI)based intervention which can filter hate content at scale. Most…

Computation and Language · Computer Science 2024-11-13 Prashant Kapil , Asif Ekbal

As a research community, we are still lacking a systematic understanding of the progress on adversarial robustness which often makes it hard to identify the most promising ideas in training robust models. A key challenge in benchmarking…

To protect users from massive hateful content, existing works studied automated hate speech detection. Despite the existing efforts, one question remains: do automated hate speech detectors conform to social media content policies? A…

Software Engineering · Computer Science 2024-03-20 Jiangrui Zheng , Xueqing Liu , Guanqun Yang , Mirazul Haque , Xing Qian , Ravishka Rathnasuriya , Wei Yang , Girish Budhrani

Hate speech detection is a critical problem in social media platforms, being often accused for enabling the spread of hatred and igniting physical violence. Hate speech detection requires overwhelming resources including high-performance…

Computation and Language · Computer Science 2020-05-14 Tomer Wullach , Amir Adler , Einat Minkov

In today's digital world, social media plays a significant role in facilitating communication and content sharing. However, the exponential rise in user-generated content has led to challenges in maintaining a respectful online environment.…

Computation and Language · Computer Science 2024-03-05 Mohammad Dehghani

Language agents increasingly act as web-enabled systems that search, browse, and synthesize information from diverse sources. However, these sources can include unreliable or adversarial content, and the robustness of agents to adversarial…

Artificial Intelligence · Computer Science 2026-03-03 Shrey Shah , Levent Ozgur

Ensuring the moral reasoning capabilities of Large Language Models (LLMs) is a growing concern as these systems are used in socially sensitive tasks. Nevertheless, current evaluation benchmarks present two major shortcomings: a lack of…

Due to the wide adoption of social media platforms like Facebook, Twitter, etc., there is an emerging need of detecting online posts that can go against the community acceptance standards. The hostility detection task has been well explored…

Computation and Language · Computer Science 2021-01-14 Arkadipta De , Venkatesh E , Kaushal Kumar Maurya , Maunendra Sankar Desarkar

Automated soft moderation systems are unable to ascertain if a post supports or refutes a false claim, resulting in a large number of contextual false positives. This limits their effectiveness, for example undermining trust in health…

Computation and Language · Computer Science 2024-07-31 Pujan Paudel , Mohammad Hammas Saeed , Rebecca Auger , Chris Wells , Gianluca Stringhini

Detecting online hate is a complex task, and low-performing models have harmful consequences when used for sensitive applications such as content moderation. Emoji-based hate is an emerging challenge for automated detection. We present…

Computation and Language · Computer Science 2022-05-09 Hannah Rose Kirk , Bertram Vidgen , Paul Röttger , Tristan Thrush , Scott A. Hale

Online texts with toxic content are a clear threat to the users on social media in particular and society in general. Although many platforms have adopted various measures (e.g., machine learning-based hate-speech detection systems) to…

Machine Learning · Computer Science 2025-04-29 Yiran Ye , Thai Le , Dongwon Lee

Well-annotated data is a prerequisite for good Natural Language Processing models. Too often, though, annotation decisions are governed by optimizing time or annotator agreement. We make a case for nuanced efforts in an interdisciplinary…

Computation and Language · Computer Science 2022-10-31 Federico Bianchi , Stefanie Anja Hills , Patricia Rossini , Dirk Hovy , Rebekah Tromble , Nava Tintarev

Despite the valuable social interactions that online media promote, these systems provide space for speech that would be potentially detrimental to different groups of people. The moderation of content imposed by many social media has…

Social and Information Networks · Computer Science 2021-08-30 Lucas Henrique Costa de Lima , Julio Reis , Philipe Melo , Fabricio Murai , Fabricio Benevenuto

Online hate is an escalating problem that negatively impacts the lives of Internet users, and is also subject to rapid changes due to evolving events, resulting in new waves of online hate that pose a critical threat. Detecting and…

Computation and Language · Computer Science 2024-07-03 Nishant Vishwamitra , Keyan Guo , Farhan Tajwar Romit , Isabelle Ondracek , Long Cheng , Ziming Zhao , Hongxin Hu

Examining the factors that the counterspeech uses are at the core of understanding the optimal methods for confronting hate speech online. Various studies have assessed the emotional base factors used in counter speech, such as emotional…

Computation and Language · Computer Science 2025-04-15 Ghadi Alyahya , Abeer Aldayel