English
Related papers

Related papers: Explainability and Hate Speech: Structured Explana…

200 papers

Content moderation faces a challenging task as social media's ability to spread hate speech contrasts with its role in promoting global connectivity. With rapidly evolving slang and hate speech, the adaptability of conventional deep…

Machine Learning · Computer Science 2024-04-18 Paras Sheth , Tharindu Kumarage , Raha Moraffah , Aman Chadha , Huan Liu

Shortcomings of current models of moderation have driven policy makers, scholars, and technologists to speculate about alternative models of content moderation. While alternative models provide hope for the future of online spaces, they can…

Computers and Society · Computer Science 2023-05-22 Sarah A. Gilbert

Online social media has become increasingly popular in recent years due to its ease of access and ability to connect with others. One of social media's main draws is its anonymity, allowing users to share their thoughts and opinions without…

Computation and Language · Computer Science 2024-04-12 Vigneshwaran Shankaran , Rajesh Sharma

Our work advances an approach for predicting hate speech in social media, drawing out the critical need to consider the discussions that follow a post to successfully detect when hateful discourse may arise. Using graph transformer…

Machine Learning · Computer Science 2023-05-02 Liam Hebert , Hong Yi Chen , Robin Cohen , Lukasz Golab

In the context of AI-based decision support systems, explanations can help users to judge when to trust the AI's suggestion, and when to question it. In this way, human oversight can prevent AI errors and biased decision-making. However,…

Human-Computer Interaction · Computer Science 2025-08-12 Laura Spillner , Rachel Ringe , Robert Porzel , Rainer Malaka

This work explores the impact of moderation on users' enjoyment of conversational AI systems. While recent advancements in Large Language Models (LLMs) have led to highly capable conversational AIs that are increasingly deployed in…

Human-Computer Interaction · Computer Science 2023-04-21 Xiaoding Lu , Aleksey Korshuk , Zongyi Liu , William Beauchamp , Chai Research

The automatic detection of hate speech online is an active research area in NLP. Most of the studies to date are based on social media datasets that contribute to the creation of hate speech detection models trained on them. However, data…

Computation and Language · Computer Science 2023-07-06 Dimosthenis Antypas , Jose Camacho-Collados

In online communities, where billions of people strive to propagate their messages, understanding how wording affects success is of primary importance. In this work, we are interested in one particularly salient aspect of wording: brevity.…

Social and Information Networks · Computer Science 2019-09-09 Kristina Gligoric , Ashton Anderson , Robert West

The prevalence of harmful content on social media platforms poses significant risks to users and society, necessitating more effective and scalable content moderation strategies. Current approaches rely on human moderators, supervised…

Computation and Language · Computer Science 2025-01-27 Akash Bonagiri , Lucen Li , Rajvardhan Oak , Zeerak Babar , Magdalena Wojcieszak , Anshuman Chhabra

This study investigates how language mutations affect the persistent diffusion of conspiracy theories on social media. Drawing on a three-year dataset of conspiracy-related posts from X, and applying computational linguistic analysis…

Computation and Language · Computer Science 2026-05-20 Calvin Yixiang Cheng , Dorian Quelle , Scott A. Hale

Social media platforms are increasingly dominated by long-form multimodal content, where harmful narratives are constructed through a complex interplay of audio, visual, and textual cues. While automated systems can flag hate speech with…

Artificial Intelligence · Computer Science 2026-05-29 Girish A. Koushik , Helen Treharne , Diptesh Kanojia

Fringe groups and organizations have a long history of using euphemisms--ordinary-sounding words with a secret meaning--to conceal what they are discussing. Nowadays, one common use of euphemisms is to evade content moderation policies…

Computation and Language · Computer Science 2021-04-01 Wanzheng Zhu , Hongyu Gong , Rohan Bansal , Zachary Weinberg , Nicolas Christin , Giulia Fanti , Suma Bhat

As large language models (LLMs) become increasingly integrated into online platforms and digital communication spaces, their potential to influence public discourse - particularly in contentious areas like climate change - requires…

Computers and Society · Computer Science 2025-06-17 Wenlu Fan , Wentao Xu

The rise of misinformation and fake news in online political discourse poses significant challenges to democratic processes and public engagement. While debunking efforts aim to counteract misinformation and foster fact-based dialogue,…

Computers and Society · Computer Science 2025-02-03 Wentao Xu , Wenlu Fan , Shiqian Lu , Tenghao Li , Bin Wang

Optimization of offensive content moderation models for different types of hateful messages is typically achieved through continued pre-training or fine-tuning on new hate speech benchmarks. However, existing benchmarks mainly address…

Computation and Language · Computer Science 2026-04-07 Irina Proskurina , Marc-Antoine Carpentier , Julien Velcin

The exponential growths of social media and micro-blogging sites not only provide platforms for empowering freedom of expressions and individual voices, but also enables people to express anti-social behaviour like online harassment,…

In this study, we tested the robustness of three communication networks extracted from the online forums included in the intranet platforms of three large companies. For each company we analyzed the communication among employees both in…

Social and Information Networks · Computer Science 2021-05-20 A. Fronzetti Colladon , F. Vagaggini

Large language models are increasingly capable of generating fluent-appearing text with relatively little task-specific supervision. But can these models accurately explain classification decisions? We consider the task of generating…

Computation and Language · Computer Science 2022-05-06 Sarah Wiegreffe , Jack Hessel , Swabha Swayamdipta , Mark Riedl , Yejin Choi

Detecting harmful content is a crucial task in the landscape of NLP applications for Social Good, with hate speech being one of its most dangerous forms. But what do we mean by hate speech, how can we define it, and how does prompting…

Computation and Language · Computer Science 2025-06-24 Matteo Melis , Gabriella Lapesa , Dennis Assenmacher

Social media platforms provide a rich environment for analyzing user behavior. Recently, deep learning-based methods have been a mainstream approach for social media analysis models involving complex patterns. However, these methods are…

Computers and Society · Computer Science 2023-08-07 Mansooreh Karami , David Mosallanezhad , Paras Sheth , Huan Liu