English
Related papers

Related papers: Digital Guardians: Can GPT-4, Perspective API, and…

200 papers

Bias in news reporting significantly impacts public perception, particularly regarding crime, politics, and societal issues. Traditional bias detection methods, predominantly reliant on human moderation, suffer from subjective…

Computation and Language · Computer Science 2025-04-07 Chen Wei Kuo , Kevin Chu , Nouar AlDahoul , Hazem Ibrahim , Talal Rahwan , Yasir Zaki

Automatic toxic language detection is critical for creating safe, inclusive online spaces. However, it is a highly subjective task, with perceptions of toxic language shaped by community norms and lived experience. Existing toxicity…

Computation and Language · Computer Science 2025-07-10 Ashima Suvarna , Christina Chance , Karolina Naranjo , Hamid Palangi , Sophie Hao , Thomas Hartvigsen , Saadia Gabriel

Online harms are a growing problem in digital spaces, putting user safety at risk and reducing trust in social media platforms. One of the most persistent forms of harm is hate speech. To address this, we need tools that combine the speed…

Computation and Language · Computer Science 2025-09-03 Paloma Piot , Diego Sánchez , Javier Parapar

This paper introduces a method for detecting inappropriately targeting language in online conversations by integrating crowd and expert annotations with ChatGPT. We focus on English conversation threads from Reddit, examining comments that…

Computation and Language · Computer Science 2025-05-23 Baran Barbarestani , Isa Maks , Piek Vossen

The proliferation of hate speech on social media platforms has necessitated the development of effective detection and moderation tools. This study evaluates the efficacy of various machine learning models in identifying hate speech and…

Computation and Language · Computer Science 2026-02-25 Saurabh Mishra , Shivani Thakur , Radhika Mamidi

Hate speech detection models are only as good as the data they are trained on. Datasets sourced from social media suffer from systematic gaps and biases, leading to unreliable models with simplistic decision boundaries. Adversarial…

Computation and Language · Computer Science 2024-03-29 Janis Goldzycher , Paul Röttger , Gerold Schneider

With the rise of voice chat rooms, a gigantic resource of data can be exposed to the research community for natural language processing tasks. Moderators in voice chat rooms actively monitor the discussions and remove the participants with…

Machine Learning · Computer Science 2021-07-13 Hadi Mansourifar , Dana Alsagheer , Reza Fathi , Weidong Shi , Lan Ni , Yan Huang

Online platforms face the challenge of moderating an ever-increasing volume of content, including harmful hate speech. In the absence of clear legal definitions and a lack of transparency regarding the role of algorithms in shaping…

Computers and Society · Computer Science 2024-06-21 David Hartmann , Amin Oueslati , Dimitri Staufer

This study evaluates the effectiveness of ChatGPT, an advanced AI model for natural language processing, in identifying targeting and inappropriate language in online comments. With the increasing challenge of moderating vast volumes of…

Computation and Language · Computer Science 2025-05-29 Barbarestani Baran , Maks Isa , Vossen Piek

Content moderation is the process of flagging content based on pre-defined platform rules. There has been a growing need for AI moderators to safeguard users as well as protect the mental health of human moderators from traumatic content.…

Computation and Language · Computer Science 2023-02-21 Meng Ye , Karan Sikka , Katherine Atwell , Sabit Hassan , Ajay Divakaran , Malihe Alikhani

This article benchmarked the ability of OpenAI's GPTs and a number of open-source LLMs to perform annotation tasks on political content. We used a novel protest event dataset comprising more than three million digital interactions and…

Computation and Language · Computer Science 2024-09-17 Bastián González-Bustamante

We present the Multi-Modal Discussion Transformer (mDT), a novel methodfor detecting hate speech in online social networks such as Reddit discussions. In contrast to traditional comment-only methods, our approach to labelling a comment as…

Computation and Language · Computer Science 2024-02-23 Liam Hebert , Gaurav Sahu , Yuxuan Guo , Nanda Kishore Sreenivas , Lukasz Golab , Robin Cohen

Since traditional social media platforms continue to ban actors spreading hate speech or other forms of abusive languages (a process known as deplatforming), these actors migrate to alternative platforms that do not moderate users content.…

Computation and Language · Computer Science 2021-11-25 Maximilian Wich , Adrian Gorniak , Tobias Eder , Daniel Bartmann , Burak Enes Çakici , Georg Groh

In this work, we demonstrate how existing classifiers for identifying toxic comments online fail to generalize to the diverse concerns of Internet users. We survey 17,280 participants to understand how user expectations for what constitutes…

Social and Information Networks · Computer Science 2021-06-09 Deepak Kumar , Patrick Gage Kelley , Sunny Consolvo , Joshua Mason , Elie Bursztein , Zakir Durumeric , Kurt Thomas , Michael Bailey

The automated detection of conspiracy theories online typically relies on supervised learning. However, creating respective training data requires expertise, time and mental resilience, given the often harmful content. Moreover, available…

Computation and Language · Computer Science 2025-01-22 Milena Pustet , Elisabeth Steffen , Helena Mihaljević

Online hate speech can harmfully impact individuals and groups, specifically on non-moderated platforms such as 4chan where users can post anonymous content. This work focuses on analysing and measuring the prevalence of online hate on…

Computation and Language · Computer Science 2025-04-02 Adrian Bermudez-Villalva , Maryam Mehrnezhad , Ehsan Toreini

Hateful speech detection is a key component of content moderation, yet current evaluation frameworks rarely assess why a text is deemed hateful. We introduce \textsf{HateXScore}, a four-component metric suite designed to evaluate the…

Computation and Language · Computer Science 2026-01-21 Yujia Hu , Roy Ka-Wei Lee

The detection of hate speech online has become an important task, as offensive language such as hurtful, obscene and insulting content can harm marginalized people or groups. This paper presents TU Berlin team experiments and results on the…

Computation and Language · Computer Science 2022-01-13 Salar Mohtaj , Vera Schmitt , Sebastian Möller

Content moderation faces a challenging task as social media's ability to spread hate speech contrasts with its role in promoting global connectivity. With rapidly evolving slang and hate speech, the adaptability of conventional deep…

Machine Learning · Computer Science 2024-04-18 Paras Sheth , Tharindu Kumarage , Raha Moraffah , Aman Chadha , Huan Liu

This study investigates the prevalence of violent language on incels.is. It evaluates GPT models (GPT-3.5 and GPT-4) for content analysis in social sciences, focusing on the impact of varying prompts and batch sizes on coding quality for…

Social and Information Networks · Computer Science 2024-01-05 Daniel Matter , Miriam Schirmer , Nir Grinberg , Jürgen Pfeffer