中文
相关论文

相关论文: A Critical Reflection on the Use of Toxicity Detec…

200 篇论文

Pretrained neural language models (LMs) are prone to generating racist, sexist, or otherwise toxic language which hinders their safe deployment. We investigate the extent to which pretrained LMs can be prompted to generate toxic language,…

计算与语言 · 计算机科学 2020-09-29 Samuel Gehman , Suchin Gururangan , Maarten Sap , Yejin Choi , Noah A. Smith

Social media platforms moderate content for each user by incorporating the outputs of both platform-wide content moderation systems and, in some cases, user-configured personal moderation preferences. However, it is unclear (1) how end…

人机交互 · 计算机科学 2025-06-13 Shagun Jhaver , Alice Qian Zhang , Quanze Chen , Nikhila Natarajan , Ruotong Wang , Amy Zhang

The challenge of automatic detection of toxic comments online has been the subject of a lot of research recently, but the focus has been mostly on detecting it in individual messages after they have been posted. Some authors have tried to…

社会与信息网络 · 计算机科学 2020-06-20 Éloi Brassard-Gourdeau , Richard Khoury

Automatic toxic language detection is critical for creating safe, inclusive online spaces. However, it is a highly subjective task, with perceptions of toxic language shaped by community norms and lived experience. Existing toxicity…

The mental health of social media users has started more and more to be put at risk by harmful, hateful, and offensive content. In this paper, we propose \textsc{StopHC}, a harmful content detection and mitigation architecture for social…

社会与信息网络 · 计算机科学 2024-11-12 Ciprian-Octavian Truică , Ana-Teodora Constantinescu , Elena-Simona Apostol

Target-group detection is the task of detecting which group(s) a piece of content is ``directed at or about''. Applications include targeted marketing, content recommendation, and group-specific content assessment. Key challenges include:…

机器学习 · 计算机科学 2026-05-05 Soumyajit Gupta , Maria De-Arteaga , Matthew Lease

Online social media platforms increasingly rely on Natural Language Processing (NLP) techniques to detect abusive content at scale in order to mitigate the harms it causes to their users. However, these techniques suffer from various…

计算与语言 · 计算机科学 2021-10-01 Sayan Ghosh , Dylan Baker , David Jurgens , Vinodkumar Prabhakaran

The widespread dissemination of hate speech, harassment, harmful and sexual content, and violence across websites and media platforms presents substantial challenges and provokes widespread concern among different sectors of society.…

计算机视觉与模式识别 · 计算机科学 2024-11-27 Nouar AlDahoul , Myles Joshua Toledo Tan , Harishwar Reddy Kasireddy , Yasir Zaki

The internet has become a central medium through which `networked publics' express their opinions and engage in debate. Offensive comments and personal attacks can inhibit participation in these spaces. Automated content moderation aims to…

计算机与社会 · 计算机科学 2017-09-06 Reuben Binns , Michael Veale , Max Van Kleek , Nigel Shadbolt

The increasing scale and complexity of online platforms raises critical policy questions around harmful content, digital well-being, and user autonomy. Traditional content moderation systems rely on centralised, top-down rules, often…

计算机与社会 · 计算机科学 2026-05-05 Ewelina Gajewska , Michal Wawer , Katarzyna Budzynska , Jaroslaw A. Chudziak

Researchers and developers increasingly rely on toxicity scoring to moderate generative language model outputs, in settings such as customer service, information retrieval, and content generation. However, toxicity scoring may render…

人机交互 · 计算机科学 2024-04-23 Jennifer Chien , Kevin R. McKee , Jackie Kay , William Isaac

The spread of toxic content on online platforms presents complex challenges that call for both theoretical insight and practical tools to test intervention strategies. In this novel research paper, we introduce a simulation-based framework…

社会与信息网络 · 计算机科学 2025-12-24 Letizia Milli , Laura Pollacci , Riccardo Guidotti

Social media platforms struggle to protect users from harmful content through content moderation. These platforms have recently leveraged machine learning models to cope with the vast amount of user-generated content daily. Since moderation…

机器学习 · 计算机科学 2023-01-27 Donghyun Son , Byounggyu Lew , Kwanghee Choi , Yongsu Baek , Seungwoo Choi , Beomjun Shin , Sungjoo Ha , Buru Chang

Toxicity has become a grave problem for many online communities and has been growing across many languages, including Russian. Hate speech creates an environment of intimidation, discrimination, and may even incite some real-world violence.…

计算与语言 · 计算机科学 2020-10-23 Nadezhda Zueva , Madina Kabirova , Pavel Kalaidin

As toxic language becomes nearly pervasive online, there has been increasing interest in leveraging the advancements in natural language processing (NLP), from very large transformer models to automatically detecting and removing toxic…

计算与语言 · 计算机科学 2020-07-02 Austin P. Wright , Omar Shaikh , Haekyu Park , Will Epperson , Muhammed Ahmed , Stephane Pinel , Diyi Yang , Duen Horng Chau

Bridging content that brings together individuals with opposing viewpoints on social media remains elusive, overshadowed by echo chambers and toxic exchanges. We propose that algorithmic curation could surface such content by considering…

社会与信息网络 · 计算机科学 2025-09-24 Ozgur Can Seckin , Bao Tran Truong , Alessandro Flammini , Filippo Menczer

With the recent rise of toxicity in online conversations on social media platforms, using modern machine learning algorithms for toxic comment detection has become a central focus of many online applications. Researchers and companies have…

人工智能 · 计算机科学 2020-03-30 Ameya Vaidya , Feng Mai , Yue Ning

With significant advances in generative AI, new technologies are rapidly being deployed with generative components. Generative models are typically trained on large datasets, resulting in model behaviors that can mimic the worst of the…

机器学习 · 计算机科学 2023-06-13 Susan Hao , Piyush Kumar , Sarah Laszlo , Shivani Poddar , Bhaktipriya Radharapu , Renee Shelby

Social media algorithms are thought to amplify variation in user beliefs, thus contributing to radicalization. However, quantitative evidence on how algorithms and user preferences jointly shape harmful online engagement is limited. I…

综合经济学 · 经济学 2025-03-11 Aarushi Kalra

Toxicity is an increasingly common and severe issue in online spaces. Consequently, a rich line of machine learning research over the past decade has focused on computationally detecting and mitigating online toxicity. These efforts…

计算与语言 · 计算机科学 2023-11-09 Wenbo Zhang , Hangzhi Guo , Ian D Kivlichan , Vinodkumar Prabhakaran , Davis Yadav , Amulya Yadav