English
Related papers

Related papers: A Critical Reflection on the Use of Toxicity Detec…

200 papers

Pretrained neural language models (LMs) are prone to generating racist, sexist, or otherwise toxic language which hinders their safe deployment. We investigate the extent to which pretrained LMs can be prompted to generate toxic language,…

Computation and Language · Computer Science 2020-09-29 Samuel Gehman , Suchin Gururangan , Maarten Sap , Yejin Choi , Noah A. Smith

Social media platforms moderate content for each user by incorporating the outputs of both platform-wide content moderation systems and, in some cases, user-configured personal moderation preferences. However, it is unclear (1) how end…

Human-Computer Interaction · Computer Science 2025-06-13 Shagun Jhaver , Alice Qian Zhang , Quanze Chen , Nikhila Natarajan , Ruotong Wang , Amy Zhang

The challenge of automatic detection of toxic comments online has been the subject of a lot of research recently, but the focus has been mostly on detecting it in individual messages after they have been posted. Some authors have tried to…

Social and Information Networks · Computer Science 2020-06-20 Éloi Brassard-Gourdeau , Richard Khoury

Automatic toxic language detection is critical for creating safe, inclusive online spaces. However, it is a highly subjective task, with perceptions of toxic language shaped by community norms and lived experience. Existing toxicity…

Computation and Language · Computer Science 2025-07-10 Ashima Suvarna , Christina Chance , Karolina Naranjo , Hamid Palangi , Sophie Hao , Thomas Hartvigsen , Saadia Gabriel

The mental health of social media users has started more and more to be put at risk by harmful, hateful, and offensive content. In this paper, we propose \textsc{StopHC}, a harmful content detection and mitigation architecture for social…

Social and Information Networks · Computer Science 2024-11-12 Ciprian-Octavian Truică , Ana-Teodora Constantinescu , Elena-Simona Apostol

Target-group detection is the task of detecting which group(s) a piece of content is ``directed at or about''. Applications include targeted marketing, content recommendation, and group-specific content assessment. Key challenges include:…

Machine Learning · Computer Science 2026-05-05 Soumyajit Gupta , Maria De-Arteaga , Matthew Lease

Online social media platforms increasingly rely on Natural Language Processing (NLP) techniques to detect abusive content at scale in order to mitigate the harms it causes to their users. However, these techniques suffer from various…

Computation and Language · Computer Science 2021-10-01 Sayan Ghosh , Dylan Baker , David Jurgens , Vinodkumar Prabhakaran

The widespread dissemination of hate speech, harassment, harmful and sexual content, and violence across websites and media platforms presents substantial challenges and provokes widespread concern among different sectors of society.…

Computer Vision and Pattern Recognition · Computer Science 2024-11-27 Nouar AlDahoul , Myles Joshua Toledo Tan , Harishwar Reddy Kasireddy , Yasir Zaki

The internet has become a central medium through which `networked publics' express their opinions and engage in debate. Offensive comments and personal attacks can inhibit participation in these spaces. Automated content moderation aims to…

Computers and Society · Computer Science 2017-09-06 Reuben Binns , Michael Veale , Max Van Kleek , Nigel Shadbolt

The increasing scale and complexity of online platforms raises critical policy questions around harmful content, digital well-being, and user autonomy. Traditional content moderation systems rely on centralised, top-down rules, often…

Computers and Society · Computer Science 2026-05-05 Ewelina Gajewska , Michal Wawer , Katarzyna Budzynska , Jaroslaw A. Chudziak

Researchers and developers increasingly rely on toxicity scoring to moderate generative language model outputs, in settings such as customer service, information retrieval, and content generation. However, toxicity scoring may render…

Human-Computer Interaction · Computer Science 2024-04-23 Jennifer Chien , Kevin R. McKee , Jackie Kay , William Isaac

The spread of toxic content on online platforms presents complex challenges that call for both theoretical insight and practical tools to test intervention strategies. In this novel research paper, we introduce a simulation-based framework…

Social and Information Networks · Computer Science 2025-12-24 Letizia Milli , Laura Pollacci , Riccardo Guidotti

Social media platforms struggle to protect users from harmful content through content moderation. These platforms have recently leveraged machine learning models to cope with the vast amount of user-generated content daily. Since moderation…

Machine Learning · Computer Science 2023-01-27 Donghyun Son , Byounggyu Lew , Kwanghee Choi , Yongsu Baek , Seungwoo Choi , Beomjun Shin , Sungjoo Ha , Buru Chang

Toxicity has become a grave problem for many online communities and has been growing across many languages, including Russian. Hate speech creates an environment of intimidation, discrimination, and may even incite some real-world violence.…

Computation and Language · Computer Science 2020-10-23 Nadezhda Zueva , Madina Kabirova , Pavel Kalaidin

As toxic language becomes nearly pervasive online, there has been increasing interest in leveraging the advancements in natural language processing (NLP), from very large transformer models to automatically detecting and removing toxic…

Computation and Language · Computer Science 2020-07-02 Austin P. Wright , Omar Shaikh , Haekyu Park , Will Epperson , Muhammed Ahmed , Stephane Pinel , Diyi Yang , Duen Horng Chau

Bridging content that brings together individuals with opposing viewpoints on social media remains elusive, overshadowed by echo chambers and toxic exchanges. We propose that algorithmic curation could surface such content by considering…

Social and Information Networks · Computer Science 2025-09-24 Ozgur Can Seckin , Bao Tran Truong , Alessandro Flammini , Filippo Menczer

With the recent rise of toxicity in online conversations on social media platforms, using modern machine learning algorithms for toxic comment detection has become a central focus of many online applications. Researchers and companies have…

Artificial Intelligence · Computer Science 2020-03-30 Ameya Vaidya , Feng Mai , Yue Ning

With significant advances in generative AI, new technologies are rapidly being deployed with generative components. Generative models are typically trained on large datasets, resulting in model behaviors that can mimic the worst of the…

Machine Learning · Computer Science 2023-06-13 Susan Hao , Piyush Kumar , Sarah Laszlo , Shivani Poddar , Bhaktipriya Radharapu , Renee Shelby

Social media algorithms are thought to amplify variation in user beliefs, thus contributing to radicalization. However, quantitative evidence on how algorithms and user preferences jointly shape harmful online engagement is limited. I…

General Economics · Economics 2025-03-11 Aarushi Kalra

Toxicity is an increasingly common and severe issue in online spaces. Consequently, a rich line of machine learning research over the past decade has focused on computationally detecting and mitigating online toxicity. These efforts…

Computation and Language · Computer Science 2023-11-09 Wenbo Zhang , Hangzhi Guo , Ian D Kivlichan , Vinodkumar Prabhakaran , Davis Yadav , Amulya Yadav