English
Related papers

Related papers: A Critical Reflection on the Use of Toxicity Detec…

200 papers

Large language models (LLMs) have exploded in popularity due to their ability to perform a wide array of natural language tasks. Text-based content moderation is one LLM use case that has received recent enthusiasm, however, there is little…

Human-Computer Interaction · Computer Science 2024-01-18 Deepak Kumar , Yousef AbuHashem , Zakir Durumeric

Malicious or manipulated prompts are known to exploit text-to-image models to generate unsafe images. Existing studies, however, focus on the passive exploitation of such harmful capabilities. In this paper, we investigate the proactive…

Cryptography and Security · Computer Science 2025-02-06 Yixin Wu , Ning Yu , Michael Backes , Yun Shen , Yang Zhang

Sensitive information detection is crucial in content moderation to maintain safe online communities. Assisting in this traditionally manual process could relieve human moderators from overwhelming and tedious tasks, allowing them to focus…

Toxicity detection mitigates the dissemination of toxic content (e.g., hateful comments, posts, and messages within online social actions) to safeguard a healthy online social environment. However, malicious users persistently develop…

Computation and Language · Computer Science 2026-03-17 Hankun Kang , Xin Miao , Jianhao Chen , Jintao Wen , Mayi Xu , Weiyu Zhang , Wenpeng Lu , Tieyun Qian

We study the impact of content moderation policies in online communities. In our theoretical model, a platform chooses a content moderation policy and individuals choose whether or not to participate in the community according to the…

Data Structures and Algorithms · Computer Science 2023-10-17 Cynthia Dwork , Chris Hays , Jon Kleinberg , Manish Raghavan

The proliferation of online hate speech has necessitated the creation of algorithms which can detect toxicity. Most of the past research focuses on this detection as a classification task, but assigning an absolute toxicity label is often…

Computation and Language · Computer Science 2022-06-28 Millon Madhur Das , Punyajoy Saha , Mithun Das

The censorship of toxic comments is often left to the judgment of imperfect models. Perspective API, a creation of Google technology incubator Jigsaw, is perhaps the most widely used toxicity classifier in industry; the model is employed by…

Computers and Society · Computer Science 2020-07-01 Elizabeth Reichert , Helen Qiu , Jasmine Bayrooti

Conversational moderation of online communities is crucial to maintaining civility for a constructive environment, but it is challenging to scale and harmful to moderators. The inclusion of sophisticated natural language generation modules…

Computation and Language · Computer Science 2024-05-07 Hyundong Cho , Shuai Liu , Taiwei Shi , Darpan Jain , Basem Rizk , Yuyang Huang , Zixun Lu , Nuan Wen , Jonathan Gratch , Emilio Ferrara , Jonathan May

Using robots in educational contexts has already shown to be beneficial for a student's learning and social behaviour. For levitating them to the next level of providing more effective and human-like tutoring, the ability to adapt to the…

Robotics · Computer Science 2022-12-21 Matthias Kraus , Diana Betancourt , Wolfgang Minker

This paper envisions a multi-agent system for detecting the presence of hate speech in online social media platforms such as Twitter and Facebook. We introduce a novel framework employing deep learning techniques to coordinate the channels…

Artificial Intelligence · Computer Science 2021-05-05 Gaurav Sahu , Robin Cohen , Olga Vechtomova

The rise of cyberbullying in social media platforms involving toxic comments has escalated the need for effective ways to monitor and moderate online interactions. Existing solutions of automated toxicity detection systems, are based on a…

Artificial Intelligence · Computer Science 2026-02-11 Michail S. Alexiou , J. Sukarno Mertoguno

The detection of hate speech or toxic content online is a complex and sensitive issue. While the identification itself is highly dependent on the context of the situation, sensitive personal attributes such as age, language, and nationality…

Multiagent Systems · Computer Science 2024-10-11 Jan Fillies , Theodoros Mitsikas , Ralph Schäfermeier , Adrian Paschke

Online communities serve as essential support channels for People Who Use Drugs (PWUD), providing access to peer support and harm reduction information. The moderation of these communities involves consequential decisions affecting member…

Human-Computer Interaction · Computer Science 2026-01-06 Kaixuan Wang , Loraine Clarke , Carl-Cyril J Dreue , Guancheng Zhou , Jason T. Jacques

Text toxicity detection systems exhibit significant biases, producing disproportionate rates of false positives on samples mentioning demographic groups. But what about toxicity detection in speech? To investigate the extent to which…

The rise of social networks has not only facilitated communication but also allowed the spread of harmful content. Although significant advances have been made in detecting toxic language in textual data, the exploration of concept-based…

Computation and Language · Computer Science 2025-12-16 Samarth Garg , Divya Singh , Deeksha Varshney , Mamta

Social media and the internet have become an integral part of how people spread and consume information. Over a period of time, social media evolved dramatically, and almost half of the population is using social media to express their…

Computation and Language · Computer Science 2021-08-03 Anjum , Rahul Katarya

Current LLMs are generally aligned to follow safety requirements and tend to refuse toxic prompts. However, LLMs can fail to refuse toxic prompts or be overcautious and refuse benign examples. In addition, state-of-the-art toxicity…

Computation and Language · Computer Science 2024-11-11 Zhanhao Hu , Julien Piet , Geng Zhao , Jiantao Jiao , David Wagner

In the pursuit of bolstering user safety, social media platforms deploy active moderation strategies, including content removal and user suspension. These measures target users engaged in discussions marked by hate speech or toxicity, often…

Social and Information Networks · Computer Science 2024-01-26 Hina Qayyum , Muhammad Ikram , Benjamin Zi Hao Zhao , Ian D. Wood , Nicolas Kourtellis , Mohamed Ali Kaafar

Content moderation is the process of screening and monitoring user-generated content online. It plays a crucial role in stopping content resulting from unacceptable behaviors such as hate speech, harassment, violence against specific…

Computation and Language · Computer Science 2023-01-02 Álvaro Huertas-García , Alejandro Martín , Javier Huertas Tato , David Camacho

Toxic and antisocial user behavior on social media platforms has received considerable scholarly attention due to its detrimental effects on society. This study takes a holistic perspective on the phenomenon of online toxicity by…

Social and Information Networks · Computer Science 2025-11-24 Lorenzo Alvisi , Victoria Popa , Guglielmo Cola , Serena Tardelli , Maurizio Tesconi
‹ Prev 1 4 5 6 7 8 10 Next ›