English
Related papers

Related papers: A Critical Reflection on the Use of Toxicity Detec…

200 papers

In health-related topics, user toxicity in online discussions frequently becomes a source of social conflict or promotion of dangerous, unscientific behaviour; common approaches for battling it include different forms of detection, flagging…

Computation and Language · Computer Science 2025-05-26 Jorge Paz-Ruza , Amparo Alonso-Betanzos , Bertha Guijarro-Berdiñas , Carlos Eiras-Franco

With the widespread use of toxic language online, platforms are increasingly using automated systems that leverage advances in natural language processing to automatically flag and remove toxic comments. However, most automated systems --…

Human-Computer Interaction · Computer Science 2021-02-11 Austin P Wright , Omar Shaikh , Haekyu Park , Will Epperson , Muhammed Ahmed , Stephane Pinel , Duen Horng Chau , Diyi Yang

Reddit administrators have generally struggled to prevent or contain such discourse for several reasons including: (1) the inability for a handful of human administrators to track and react to millions of posts and comments per day and (2)…

Social and Information Networks · Computer Science 2019-07-01 Hussam Habib , Maaz Bin Musa , Fareed Zaffar , Rishab Nithyanand

Existing toxic detection models face significant limitations, such as lack of transparency, customization, and reproducibility. These challenges stem from the closed-source nature of their training data and the paucity of explanations for…

Computation and Language · Computer Science 2025-01-24 Tinh Son Luong , Thanh-Thien Le , Thang Viet Doan , Linh Ngo Van , Thien Huu Nguyen , Diep Thi-Ngoc Nguyen

Online communities have gained considerable importance in recent years due to the increasing number of people connected to the Internet. Moderating user content in online communities is mainly performed manually, and reducing the workload…

Information Retrieval · Computer Science 2019-01-16 Etienne Papegnies , Vincent Labatut , Richard Dufour , Georges Linares

Incivility remains a major challenge for online discussion platforms, to such an extent that even conversations between well-intentioned users can often derail into uncivil behavior. Traditionally, platforms have relied on moderators to --…

Human-Computer Interaction · Computer Science 2022-12-06 Jonathan P. Chang , Charlotte Schluger , Cristian Danescu-Niculescu-Mizil

Toxicity and abuse are common in online peer-production communities. The social structure of peer-production communities that aim to produce accurate and trustworthy information require some conflict and gate-keeping to spur content…

Human-Computer Interaction · Computer Science 2023-03-27 Chris Blakely , Andrew Vargo

Lack of moderation in online communities enables participants to incur in personal aggression, harassment or cyberbullying, issues that have been accentuated by extremist radicalisation in the contemporary post-truth politics scenario. This…

Computation and Language · Computer Science 2018-01-08 Nestor Rodriguez , Sergio Rojas-Galeano

Social media platforms increasingly employ proactive moderation techniques, such as detecting and curbing toxic and uncivil comments, to prevent the spread of harmful content. Despite these efforts, such approaches are often criticized for…

Human-Computer Interaction · Computer Science 2025-07-30 Xiaotian Su , Naim Zierau , Soomin Kim , April Yi Wang , Thiemo Wambsganss

The detrimental effects of toxicity in competitive online video games are widely acknowledged, prompting publishers to monitor player chat conversations. This is challenging due to the context-dependent nature of toxicity, often spread…

Computation and Language · Computer Science 2025-04-03 Adrien Schurger-Foy , Rafal Dariusz Kocielnik , Caglar Gulcehre , R. Michael Alvarez

The proliferation of social media platforms and online communities has inadvertently catalyzed the spread of cyberbullying, hate speech, and other forms of online toxicity, making the effective governance of such harm a critical societal…

Artificial Intelligence · Computer Science 2026-05-28 Yiting Huang , Wenting Zhu , Zekun Wang , Qingpo Yang , Yakai Chen , Zihui Xu , Yueyue Zhang , Sanchuan Guo , Xi Zhang

Extensive efforts in automated approaches for content moderation have been focused on developing models to identify toxic, offensive, and hateful content with the aim of lightening the load for moderators. Yet, it remains uncertain whether…

Computation and Language · Computer Science 2024-11-14 Yang Trista Cao , Lovely-Frances Domingo , Sarah Ann Gilbert , Michelle Mazurek , Katie Shilton , Hal Daumé

In this work, we demonstrate how existing classifiers for identifying toxic comments online fail to generalize to the diverse concerns of Internet users. We survey 17,280 participants to understand how user expectations for what constitutes…

Social and Information Networks · Computer Science 2021-06-09 Deepak Kumar , Patrick Gage Kelley , Sunny Consolvo , Joshua Mason , Elie Bursztein , Zakir Durumeric , Kurt Thomas , Michael Bailey

Although there have been automated approaches and tools supporting toxicity censorship for social posts, most of them focus on detection. Toxicity censorship is a complex process, wherein detection is just an initial task and a user can…

Human-Computer Interaction · Computer Science 2025-05-23 Yaqiong Li , Peng Zhang , Hansu Gu , Tun Lu , Siyuan Qiao , Yubo Shu , Yiyang Shao , Ning Gu

Despite remarkable advances that large language models have achieved in chatbots, maintaining a non-toxic user-AI interactive environment has become increasingly critical nowadays. However, previous efforts in toxicity detection have been…

Computation and Language · Computer Science 2023-10-27 Zi Lin , Zihan Wang , Yongqi Tong , Yangkun Wang , Yuxin Guo , Yujia Wang , Jingbo Shang

Large Language Models are widely used for content moderation but often present certain over-sensitivity, leading to misclassification of benign content and rejecting safe user commands. While previous research attributes this issue…

Computation and Language · Computer Science 2026-03-19 Yuxin Wang , Botao Yu , Ivory Yang , Saeed Hassanpour , Soroush Vosoughi

The detection of sensitive content in large datasets is crucial for ensuring that shared and analysed data is free from harmful material. However, current moderation tools, such as external APIs, suffer from limitations in customisation,…

Computation and Language · Computer Science 2025-06-25 Dimosthenis Antypas , Indira Sen , Carla Perez-Almendros , Jose Camacho-Collados , Francesco Barbieri

The rapid growth of live-streaming platforms such as Twitch has introduced complex challenges in moderating toxic behavior. Traditional moderation approaches, such as human annotation and keyword-based filtering, have demonstrated utility,…

Computation and Language · Computer Science 2026-02-05 Baktash Ansari , Elias Martin , Afra Mashhadi

Although automated harmful content detection systems are frequently used to monitor online platforms, moderators and end users frequently cannot understand the logic underlying their predictions. While recent studies have focused on…

Computation and Language · Computer Science 2026-03-20 Trishita Dhara , Siddhesh Sheth

Moderation is crucial to promoting healthy on-line discussions. Although several `toxicity' detection datasets and models have been published, most of them ignore the context of the posts, implicitly assuming that comments maybe judged…

Computation and Language · Computer Science 2020-06-02 John Pavlopoulos , Jeffrey Sorensen , Lucas Dixon , Nithum Thain , Ion Androutsopoulos