中文
相关论文

相关论文: What If Moderation Didn't Mean Suppression? A Case…

200 篇论文

Polarization is implicated in the erosion of democracy and the progression to violence, which makes the polarization properties of large algorithmic content selection systems (recommender systems) a matter of concern for peace and security.…

信息检索 · 计算机科学 2021-07-13 Jonathan Stray

When toxic online communities on mainstream platforms face moderation measures, such as bans, they may migrate to other platforms with laxer policies or set up their own dedicated websites. Previous work suggests that within mainstream…

Content moderation research typically prioritizes representing and addressing challenges for one group of stakeholders or communities in one type of context. While taking a focused approach is reasonable or even favorable for empirical case…

人机交互 · 计算机科学 2022-06-08 Jialun Aaron Jiang , Peipei Nie , Jed R. Brubaker , Casey Fiesler

Accurately estimating how users respond to moderation interventions is paramount for developing effective and user-centred moderation strategies. However, this requires a clear understanding of which user characteristics are associated with…

计算机与社会 · 计算机科学 2025-10-24 Benedetta Tessa , Alejandro Moreo , Stefano Cresci , Tiziano Fagni , Fabrizio Sebastiani

Generative AI systems powered by Large Language Models (LLMs) usually use content moderation to prevent harmful content spread. To evaluate the robustness of content moderation, several metamorphic testing techniques have been proposed to…

软件工程 · 计算机科学 2025-03-24 Honghao Tan , Haibo Wang , Diany Pressato , Yisen Xu , Shin Hwei Tan

The emergence of multi-agent systems introduces novel moderation challenges that extend beyond content filtering. Agents with malicious intent may contribute harmful content that appears benign to evade content-based moderation, while…

人工智能 · 计算机科学 2026-05-15 Ali Al-Lawati , Nafis Tripto , Abolfazl Ansari , Jason Lucas , Suhang Wang , Dongwon Lee

Hate speech remains a pressing challenge on social media, where platform moderation often fails to protect targeted users. Personal moderation tools that let users decide how content is filtered can address some of these shortcomings.…

人机交互 · 计算机科学 2026-03-03 Anna Ricarda Luther , Hendrik Heuer , Stephanie Geise , Sebastian Haunss , Andreas Breiter

Content moderation is the process of flagging content based on pre-defined platform rules. There has been a growing need for AI moderators to safeguard users as well as protect the mental health of human moderators from traumatic content.…

计算与语言 · 计算机科学 2023-02-21 Meng Ye , Karan Sikka , Katherine Atwell , Sabit Hassan , Ajay Divakaran , Malihe Alikhani

With the growth of social media and large language models, content moderation has become crucial. Many existing datasets lack adequate representation of different groups, resulting in unreliable assessments. To tackle this, we propose a…

计算与语言 · 计算机科学 2024-12-19 Shanu Kumar , Gauri Kholkar , Saish Mendke , Anubhav Sadana , Parag Agrawal , Sandipan Dandapat

This paper aims to survey various techniques utilized for content moderation in end-to-end encryption systems. We assess the challenging aspect of content moderation: maintaining a safe platform while assuring user privacy. We study the…

密码学与安全 · 计算机科学 2025-01-31 Chaitanya Rahalkar , Anushka Virgaonkar

The widespread dissemination of hate speech, harassment, harmful and sexual content, and violence across websites and media platforms presents substantial challenges and provokes widespread concern among different sectors of society.…

计算机视觉与模式识别 · 计算机科学 2024-11-27 Nouar AlDahoul , Myles Joshua Toledo Tan , Harishwar Reddy Kasireddy , Yasir Zaki

There is an ongoing debate about how to moderate toxic speech on social media and the impact of content moderation on online discourse. This paper proposes and validates a methodology for measuring the content-moderation-induced distortions…

社会与信息网络 · 计算机科学 2026-03-04 Mahyar Habibi , Dirk Hovy , Carlo Schwarz

Large language models (LLMs) are increasingly used in content moderation systems, where ensuring fairness and neutrality is essential. In this study, we examine how persona adoption influences the consistency and fairness of harmful content…

计算与语言 · 计算机科学 2025-10-31 Stefano Civelli , Pietro Bernardelle , Nardiena A. Pratama , Gianluca Demartini

Social media has shaken the foundations of our society, unlikely as it may seem. Many of the popular tools used to moderate harmful digital content, however, have received widespread criticism from both the academic community and the public…

社会与信息网络 · 计算机科学 2020-10-21 Jiachen Jiang , Soroush Vosoughi

Online video platforms receive hundreds of hours of uploads every minute, making manual content moderation impossible. Unfortunately, the most vulnerable consumers of malicious video content are children from ages 1-5 whose attention is…

计算机视觉与模式识别 · 计算机科学 2023-05-26 Syed Hammad Ahmed , Muhammad Junaid Khan , H. M. Umer Qaisar , Gita Sukthankar

Though detection systems have been developed to identify obscene content such as pornography and violence, artificial intelligence is simply not good enough to fully automate this task yet. Due to the need for manual verification, social…

人机交互 · 计算机科学 2020-01-07 Brandon Dang , Martin J. Riedl , Matthew Lease

We present Moderator, a policy-based model management system that allows administrators to specify fine-grained content moderation policies and modify the weights of a text-to-image (TTI) model to make it significantly more challenging for…

密码学与安全 · 计算机科学 2024-09-13 Peiran Wang , Qiyu Li , Longxuan Yu , Ziyao Wang , Ang Li , Haojian Jin

Personal moderation tools on social media platforms let users control their news feeds by configuring acceptable toxicity thresholds for their feed content or muting inappropriate accounts. This research examines how four critical…

人机交互 · 计算机科学 2024-03-05 Shagun Jhaver

Online content moderation is essential for maintaining a healthy digital environment, and reliance on AI for this task continues to grow. Consider a user comment using national stereotypes to insult a politician. This example illustrates…

人工智能 · 计算机科学 2026-03-03 Houde Dong , Yifei She , Kai Ye , Liangcai Su , Chenxiong Qian , Jie Hao

The detection of hate speech or toxic content online is a complex and sensitive issue. While the identification itself is highly dependent on the context of the situation, sensitive personal attributes such as age, language, and nationality…

多智能体系统 · 计算机科学 2024-10-11 Jan Fillies , Theodoros Mitsikas , Ralph Schäfermeier , Adrian Paschke