中文
相关论文

相关论文: A Critical Reflection on the Use of Toxicity Detec…

200 篇论文

Today's large-scale algorithmic and automated deployment of decision-making systems threatens to exclude marginalized communities. Thus, the emergent danger comes from the effectiveness and the propensity of such systems to replicate,…

计算机与社会 · 计算机科学 2022-09-13 Kristine Gloria , Nidhi Rastogi , Stevie DeGroff

We present Moderator, a policy-based model management system that allows administrators to specify fine-grained content moderation policies and modify the weights of a text-to-image (TTI) model to make it significantly more challenging for…

密码学与安全 · 计算机科学 2024-09-13 Peiran Wang , Qiyu Li , Longxuan Yu , Ziyao Wang , Ang Li , Haojian Jin

Content moderation is a widely used strategy to prevent the dissemination of irregular information on social media platforms. Despite extensive research on developing automated models to support decision-making in content moderation, there…

社会与信息网络 · 计算机科学 2024-08-23 Wangjiaxuan Xin , Kanlun Wang , Zhe Fu , Lina Zhou

Interpretations of a single sentence can vary, particularly when its context is lost. This paper aims to simulate how readers perceive content with varying toxicity levels by generating diverse interpretations of out-of-context sentences.…

计算与语言 · 计算机科学 2026-04-17 Maria Mihaela Trusca , Liesbeth Allein

Online multiplayer games like League of Legends, Counter Strike, and Skribbl.io create experiences through community interactions. Providing players with the ability to interact with each other through multiple modes also opens a Pandora…

计算机与社会 · 计算机科学 2024-07-08 Lynnette Hui Xian Ng , Adrian Xuan Wei Lim , Michael Miller Yoder

The pervasiveness of online toxicity, including hate speech and trolling, disrupts digital interactions and online well-being. Previous research has mainly focused on post-hoc moderation, overlooking the real-time emotional dynamics of…

人机交互 · 计算机科学 2025-09-03 Akriti Verma , Shama Islam , Valeh Moghaddam , Adnan Anwar

WhatsApp groups have become a hotbed for the propagation of harmful content including misinformation, hate speech, polarizing content, and rumors, especially in Global South countries. Given the platform's end-to-end encryption, moderation…

人机交互 · 计算机科学 2024-08-20 Dhruv Agarwal , Farhana Shahid , Aditya Vashistha

Risk perception is subjective, and youth's understanding of toxic content differs from that of adults. Although previous research has conducted extensive studies on toxicity detection in social media, the investigation of youth's unique…

计算与语言 · 计算机科学 2025-08-05 Yaqiong Li , Peng Zhang , Lin Wang , Hansu Gu , Siyuan Qiao , Ning Gu , Tun Lu

Text classification methods have been widely investigated as a way to detect content of low credibility: fake news, social media bots, propaganda, etc. Quite accurate models (likely based on deep neural networks) help in moderating public…

计算与语言 · 计算机科学 2026-03-04 Piotr Przybyła , Alexander Shvets , Horacio Saggion

While algorithm audits are growing rapidly in commonality and public importance, relatively little scholarly work has gone toward synthesizing prior work and strategizing future research in the area. This systematic literature review aims…

计算机与社会 · 计算机科学 2021-02-09 Jack Bandy

As Machine Learning models continue to be relied upon for making automated decisions, the issue of model bias becomes more and more prevalent. In this paper, we approach training a text classifica-tion model and optimize on bias…

计算与语言 · 计算机科学 2019-08-19 Apik Ashod Zorian , Chandra Shekar Bikkanur

Warning: this paper contains content that maybe offensive or upsetting. Recent research in Natural Language Processing (NLP) has advanced the development of various toxicity detection models with the intention of identifying and mitigating…

计算与语言 · 计算机科学 2022-05-06 Ninareh Mehrabi , Ahmad Beirami , Fred Morstatter , Aram Galstyan

Content moderation research typically prioritizes representing and addressing challenges for one group of stakeholders or communities in one type of context. While taking a focused approach is reasonable or even favorable for empirical case…

人机交互 · 计算机科学 2022-06-08 Jialun Aaron Jiang , Peipei Nie , Jed R. Brubaker , Casey Fiesler

The proliferation of social media platforms has led to an increase in the spread of hate speech, particularly targeting vulnerable communities. Unfortunately, existing methods for automatically identifying and blocking toxic language rely…

计算与语言 · 计算机科学 2025-02-24 Shiza Ali , Jeremy Blackburn , Gianluca Stringhini

Open-source large language models are becoming increasingly available and popular among researchers and practitioners. While significant progress has been made on open-weight models, open training data is a practice yet to be adopted by the…

计算与语言 · 计算机科学 2024-11-19 Catherine Arnett , Eliot Jones , Ivan P. Yamshchikov , Pierre-Carl Langlais

When trained on large, unfiltered crawls from the internet, language models pick up and reproduce all kinds of undesirable biases that can be found in the data: they often generate racist, sexist, violent or otherwise toxic language. As…

计算与语言 · 计算机科学 2021-09-10 Timo Schick , Sahana Udupa , Hinrich Schütze

Large Language Models have demonstrated impressive fluency across diverse tasks, yet their tendency to produce toxic content remains a critical challenge for AI safety and public trust. Existing toxicity mitigation approaches primarily…

计算与语言 · 计算机科学 2025-09-23 Zuhair Hasan Shaik , Abdullah Mazhar , Aseem Srivastava , Md Shad Akhtar

As the last few years have seen an increase in online hostility and polarization both, we need to move beyond the fack-checking reflex or the praise for better moderation on social networking sites (SNS) and investigate their impact on…

离散数学 · 计算机科学 2023-03-28 David Chavalarias , Paul Bouchaud , Maziyar Panahi

This paper aims to help structure the risk landscape associated with large-scale Language Models (LMs). In order to foster advances in responsible innovation, an in-depth understanding of the potential risks posed by these models is needed.…

Social media platforms have been establishing content moderation guidelines and employing various moderation policies to counter hate speech and misinformation. The goal of this paper is to study these community guidelines and moderation…

计算机与社会 · 计算机科学 2023-08-10 Mohit Singhal , Chen Ling , Pujan Paudel , Poojitha Thota , Nihal Kumarswamy , Gianluca Stringhini , Shirin Nilizadeh
‹ 上一页 1 8 9 10 下一页 ›