English
Related papers

Related papers: A Critical Reflection on the Use of Toxicity Detec…

200 papers

The emergence of multi-agent systems introduces novel moderation challenges that extend beyond content filtering. Agents with malicious intent may contribute harmful content that appears benign to evade content-based moderation, while…

Artificial Intelligence · Computer Science 2026-05-15 Ali Al-Lawati , Nafis Tripto , Abolfazl Ansari , Jason Lucas , Suhang Wang , Dongwon Lee

The rise of hate speech on online platforms has led to an urgent need for effective content moderation. However, the subjective and multi-faceted nature of hateful online content, including implicit hate speech, poses significant challenges…

Computation and Language · Computer Science 2023-03-17 Uma Gunturi , Xiaohan Ding , Eugenia H. Rho

Harmful content detection models tend to have higher false positive rates for content from marginalized groups. In the context of marginal abuse modeling on Twitter, such disproportionate penalization poses the risk of reduced visibility,…

Computation and Language · Computer Science 2022-10-13 Kyra Yee , Alice Schoenauer Sebag , Olivia Redfield , Emily Sheng , Matthias Eck , Luca Belli

Real-time toxicity detection in online environments poses a significant challenge, due to the increasing prevalence of social media and gaming platforms. We introduce ToxBuster, a simple and scalable model that reliably detects toxic…

Computation and Language · Computer Science 2024-08-22 Zachary Yang , Nicolas Grenan-Godbout , Reihaneh Rabbany

Now-a-days, derogatory comments are often made by one another, not only in offline environment but also immensely in online environments like social networking websites and online communities. So, an Identification combined with Prevention…

Computation and Language · Computer Science 2019-03-19 Navoneel Chakrabarty

Research on children's online experience and computer interaction often overlooks the relationship children have with hidden algorithms that control the content they encounter. Furthermore, it is not only about how children interact with…

Human-Computer Interaction · Computer Science 2024-06-13 Belén Saldías

Machine learning (ML) is widely used to moderate online content. Despite its scalability relative to human moderation, the use of ML introduces unique challenges to content moderation. One such challenge is predictive multiplicity: multiple…

Computers and Society · Computer Science 2024-02-28 Juan Felipe Gomez , Caio Vieira Machado , Lucas Monteiro Paes , Flavio P. Calmon

Social media platforms provide an environment where people can freely engage in discussions. Unfortunately, they also enable several problems, such as online harassment. Recently, Google and Jigsaw started a project called Perspective,…

Machine Learning · Computer Science 2017-02-28 Hossein Hosseini , Sreeram Kannan , Baosen Zhang , Radha Poovendran

Online social media has become increasingly popular in recent years due to its ease of access and ability to connect with others. One of social media's main draws is its anonymity, allowing users to share their thoughts and opinions without…

Computation and Language · Computer Science 2024-04-12 Vigneshwaran Shankaran , Rajesh Sharma

Toxic language detection systems often falsely flag text that contains minority group mentions as toxic, as those groups are often the targets of online hate. Such over-reliance on spurious correlations also causes systems to struggle with…

Computation and Language · Computer Science 2022-07-15 Thomas Hartvigsen , Saadia Gabriel , Hamid Palangi , Maarten Sap , Dipankar Ray , Ece Kamar

While recent research has focused on developing safeguards for generative AI (GAI) model-level content safety, little is known about how content moderation to prevent malicious content performs for end-users in real-world GAI products. To…

Human-Computer Interaction · Computer Science 2025-06-18 Lan Gao , Oscar Chen , Rachel Lee , Nick Feamster , Chenhao Tan , Marshini Chetty

Now that AI-driven moderation has become pervasive in everyday life, we often hear claims that "the AI is biased". While this is often said jokingly, the light-hearted remark reflects a deeper concern. How can we be certain that an online…

Computation and Language · Computer Science 2026-04-02 Subhojit Ghimire

Toxicity detection is crucial for maintaining the peace of the society. While existing methods perform well on normal toxic contents or those generated by specific perturbation methods, they are vulnerable to evolving perturbation patterns.…

Cryptography and Security · Computer Science 2025-03-05 Hankun Kang , Jianhao Chen , Yongqi Li , Xin Miao , Mayi Xu , Ming Zhong , Yuanyuan Zhu , Tieyun Qian

Online platforms and communities establish their own norms that govern what behavior is acceptable within the community. Substantial effort in NLP has focused on identifying unacceptable behaviors and, recently, on forecasting them before…

Computation and Language · Computer Science 2021-10-12 Chan Young Park , Julia Mendelsohn , Karthik Radhakrishnan , Kinjal Jain , Tushar Kanakagiri , David Jurgens , Yulia Tsvetkov

With surge in online platforms, there has been an upsurge in the user engagement on these platforms via comments and reactions. A large portion of such textual comments are abusive, rude and offensive to the audience. With machine learning…

Computation and Language · Computer Science 2021-08-17 Ayush Kumar , Pratik Kumar

Content moderation and toxicity classification represent critical tasks with significant social implications. However, studies have shown that major classification models exhibit tendencies to magnify or reduce biases and potentially…

Computation and Language · Computer Science 2024-11-28 Haniyeh Ehsani Oskouie , Christina Chance , Claire Huang , Margaret Capetz , Elizabeth Eyeson , Majid Sarrafzadeh

Effective content moderation systems require explicit classification criteria, yet online communities like subreddits often operate with diverse, implicit standards. This work introduces a novel approach to identify and extract these…

Computation and Language · Computer Science 2025-09-04 Youngwoo Kim , Himanshu Beniwal , Steven L. Johnson , Thomas Hartvigsen

Hate speech on online platforms has been credibly linked to multiple instances of real world violence. This calls for an urgent need to understand how toxic content spreads and how it might be mitigated on online social networks, and…

Social and Information Networks · Computer Science 2025-11-26 Aatman Vaidya , Harsh Bhagat , Seema Nagar , Amit A. Nanavati

Existing studies have investigated the tendency of autoregressive language models to generate contexts that exhibit undesired biases and toxicity. Various debiasing approaches have been proposed, which are primarily categorized into…

Computation and Language · Computer Science 2022-05-03 Yoon A Park , Frank Rudzicz

Online platforms take proactive measures to detect and address undesirable behavior, aiming to focus these resource-intensive efforts where such behavior is most prevalent. This article considers the problem of efficient sampling for…

Machine Learning · Computer Science 2025-03-28 Jacob Morrier , Rafal Kocielnik , R. Michael Alvarez