English
Related papers

Related papers: Detecting Inappropriate Messages on Sensitive Topi…

200 papers

Online platforms and communities establish their own norms that govern what behavior is acceptable within the community. Substantial effort in NLP has focused on identifying unacceptable behaviors and, recently, on forecasting them before…

Computation and Language · Computer Science 2021-10-12 Chan Young Park , Julia Mendelsohn , Karthik Radhakrishnan , Kinjal Jain , Tushar Kanakagiri , David Jurgens , Yulia Tsvetkov

Large language models produce human-like text that drive a growing number of applications. However, recent literature and, increasingly, real world observations, have demonstrated that these models can generate language that is toxic,…

Large Language Models are widely used for content moderation but often present certain over-sensitivity, leading to misclassification of benign content and rejecting safe user commands. While previous research attributes this issue…

Computation and Language · Computer Science 2026-03-19 Yuxin Wang , Botao Yu , Ivory Yang , Saeed Hassanpour , Soroush Vosoughi

Bridging content that brings together individuals with opposing viewpoints on social media remains elusive, overshadowed by echo chambers and toxic exchanges. We propose that algorithmic curation could surface such content by considering…

Social and Information Networks · Computer Science 2025-09-24 Ozgur Can Seckin , Bao Tran Truong , Alessandro Flammini , Filippo Menczer

Online social media has become increasingly popular in recent years due to its ease of access and ability to connect with others. One of social media's main draws is its anonymity, allowing users to share their thoughts and opinions without…

Computation and Language · Computer Science 2024-04-12 Vigneshwaran Shankaran , Rajesh Sharma

Personalized AI systems, from recommendation systems to chatbots, are a prevalent method for distributing content to users based on their learned preferences. However, there is growing concern about the adverse effects of these systems,…

Information Retrieval · Computer Science 2025-09-10 Amelia Kovacs , Jerry Chee , Kimia Kazemian , Sarah Dean

The challenge of automatic detection of toxic comments online has been the subject of a lot of research recently, but the focus has been mostly on detecting it in individual messages after they have been posted. Some authors have tried to…

Social and Information Networks · Computer Science 2020-06-20 Éloi Brassard-Gourdeau , Richard Khoury

Toxic language includes content that is offensive, abusive, or that promotes harm. Progress in preventing toxic output from large language models (LLMs) is hampered by inconsistent definitions of toxicity. We introduce TRuST, a large-scale…

Computation and Language · Computer Science 2026-01-07 Berk Atil , Namrata Sureddy , Rebecca J. Passonneau

Tackling toxic behavior in digital communication continues to be a pressing concern for both academics and industry professionals. While significant research has explored toxicity on platforms like social networks and discussion boards,…

Computation and Language · Computer Science 2025-09-01 Naquee Rizwan , Nayandeep Deb , Sarthak Roy , Vishwajeet Singh Solanki , Kiran Garimella , Animesh Mukherjee

The prevalence and impact of toxic discussions online have made content moderation crucial.Automated systems can play a vital role in identifying toxicity, and reducing the reliance on human moderation.Nevertheless, identifying toxic…

Artificial Intelligence · Computer Science 2023-11-02 Senjuti Dutta , Sid Mittal , Sherol Chen , Deepak Ramachandran , Ravi Rajakumar , Ian Kivlichan , Sunny Mak , Alena Butryna , Praveen Paritosh

Collecting annotations from human raters often results in a trade-off between the quantity of labels one wishes to gather and the quality of these labels. As such, it is often only possible to gather a small amount of high-quality labels.…

Machine Learning · Computer Science 2021-10-05 Neel Nanda , Jonathan Uesato , Sven Gowal

Detecting toxic language including sexism, harassment and abusive behaviour, remains a critical challenge, particularly in its subtle and context-dependent forms. Existing approaches largely focus on isolated message-level classification,…

Toxic language is difficult to define, as it is not monolithic and has many variations in perceptions of toxicity. This challenge of detecting toxic language is increased by the highly contextual and subjectivity of its interpretation,…

Computation and Language · Computer Science 2023-05-19 Huriyyah Althunayan , Rahaf Bahlas , Manar Alharbi , Lena Alsuwailem , Abeer Aldayel , Rehab ALahmadi

Toxic sentiment analysis on Twitter (X) often focuses on specific topics and events such as politics and elections. Datasets of toxic users in such research are typically gathered through lexicon-based techniques, providing only a…

Social and Information Networks · Computer Science 2024-06-06 Hina Qayyum , Muhammad Ikram , Benjamin Zhao , Ian Wood , Mohamad Ali Kaafar , Nicolas Kourtellis

Moderation is crucial to promoting healthy on-line discussions. Although several `toxicity' detection datasets and models have been published, most of them ignore the context of the posts, implicitly assuming that comments maybe judged…

Computation and Language · Computer Science 2020-06-02 John Pavlopoulos , Jeffrey Sorensen , Lucas Dixon , Nithum Thain , Ion Androutsopoulos

A lack of demographic context in existing toxic speech datasets limits our understanding of how different age groups communicate online. In collaboration with funk, a German public service content network, this research introduces the first…

Computation and Language · Computer Science 2025-09-01 Jan Fillies , Michael Peter Hoffmann , Rebecca Reichel , Roman Salzwedel , Sven Bodemer , Adrian Paschke

Twitter is one of the most popular online micro-blogging and social networking platforms. This platform allows individuals to freely express opinions and interact with others regardless of geographic barriers. However, with the good that…

Social and Information Networks · Computer Science 2022-11-09 Nazanin Salehabadi , Anne Groggel , Mohit Singhal , Sayak Saha Roy , Shirin Nilizadeh

Many under-resourced languages require high-quality datasets for specific tasks such as offensive language detection, disinformation, or misinformation identification. However, the intricacies of the content may have a detrimental effect on…

Computation and Language · Computer Science 2023-11-20 Stetsenko Daria

Toxicity detection algorithms, originally designed with reactive content moderation in mind, are increasingly being deployed into proactive end-user interventions to moderate content. Through a socio-technical lens and focusing on contexts…

Human-Computer Interaction · Computer Science 2025-02-25 Mark Warner , Angelika Strohmayer , Matthew Higgs , Lynne Coventry

Large language models (LLMs) have become integral to various real-world applications, leveraging massive, web-sourced datasets like Common Crawl, C4, and FineWeb for pretraining. While these datasets provide linguistic data essential for…

Computation and Language · Computer Science 2025-08-14 Sai Krishna Mendu , Harish Yenala , Aditi Gulati , Shanu Kumar , Parag Agrawal