中文
相关论文

相关论文: A Trade-off-centered Framework of Content Moderati…

200 篇论文

Globally distributed groups require collaborative systems to support their work. Besides being able to support the teamwork, these systems also should promote well-being and maximize the human potential that leads to an engaging system and…

人机交互 · 计算机科学 2019-04-09 Irawan Nurhas , Jan Pawlowski , Stefan Geisler , Maria Kovtunenko , Bayu Rima Aditya

To address the widespread problem of uncivil behavior, many online discussion platforms employ human moderators to take action against objectionable content, such as removing it or placing sanctions on its authors. This reactive paradigm of…

计算机与社会 · 计算机科学 2022-12-01 Charlotte Schluger , Jonathan P. Chang , Cristian Danescu-Niculescu-Mizil , Karen Levy

Decentralising the Web is a desirable but challenging goal. One particular challenge is achieving decentralised content moderation in the face of various adversaries (e.g. trolls). To overcome this challenge, many Decentralised Web (DW)…

Online communities rely on a mix of platform policies and community-authored rules to define acceptable behavior and maintain order. However, these rules vary widely across communities, evolve over time, and are enforced inconsistently,…

计算机与社会 · 计算机科学 2025-10-09 Mattia Samory , Diana Pamfile , Andrew To , Shruti Phadke

The sheer volume of online user-generated content has rendered content moderation technologies essential in order to protect digital platform audiences from content that may cause anxiety, worry, or concern. Despite the efforts towards…

计算机视觉与模式识别 · 计算机科学 2022-12-02 Ioannis Sarridis , Christos Koutlis , Olga Papadopoulou , Symeon Papadopoulos

Understanding the bias-variance tradeoff in user representation learning is essential for improving recommendation quality in modern content platforms. While well studied in static settings, this tradeoff becomes significantly more complex…

计算机科学与博弈论 · 计算机科学 2026-03-03 Kang Wang , Renzhe Xu , Bo Li

This paper investigates how content moderation affects content creation in an ideologically diverse online environments. We develop a model in which users act as both creators and consumers, differing in their ideological affiliation and…

综合经济学 · 经济学 2025-11-26 Ying Bao , Jessie Liu

Extensive efforts in automated approaches for content moderation have been focused on developing models to identify toxic, offensive, and hateful content with the aim of lightening the load for moderators. Yet, it remains uncertain whether…

计算与语言 · 计算机科学 2024-11-14 Yang Trista Cao , Lovely-Frances Domingo , Sarah Ann Gilbert , Michelle Mazurek , Katie Shilton , Hal Daumé

High-stakes applications rely on combining Artificial Intelligence (AI) and humans for responsive and reliable decision making. For example, content moderation in social media platforms often employs an AI-human pipeline to promptly remove…

机器学习 · 计算机科学 2025-08-14 Thodoris Lykouris , Wentao Weng

This chapter focuses on the intersection of user experience (UX) and wellbeing in the context of content moderation. Human content moderators play a key role in protecting end users from harm by detecting, evaluating, and addressing content…

人机交互 · 计算机科学 2025-09-10 Diana Mihalache , Dalila Szostak

Online abuse is becoming an increasingly prevalent issue in modern-day society, with 41 percent of Americans having experienced online harassment in some capacity in 2021. People who identify as women, in particular, can be subjected to a…

人机交互 · 计算机科学 2023-01-19 Sarah Barrington

Large language models and LLM-based agents are increasingly used for cybersecurity tasks that are inherently dual-use. Existing approaches to refusal, spanning academic policy frameworks and commercially deployed systems, often rely on…

计算与语言 · 计算机科学 2026-02-19 Noa Linder , Meirav Segal , Omer Antverg , Gil Gekker , Tomer Fichman , Omri Bodenheimer , Edan Maor , Omer Nevo

In the evolving landscape of online communication, moderating hate speech (HS) presents an intricate challenge, compounded by the multimodal nature of digital content. This comprehensive survey delves into the recent strides in HS…

计算与语言 · 计算机科学 2024-10-31 Ming Shan Hee , Shivam Sharma , Rui Cao , Palash Nandi , Preslav Nakov , Tanmoy Chakraborty , Roy Ka-Wei Lee

Content moderation at scale remains one of the most pressing challenges in today's digital ecosystem, where billions of user- and AI-generated artifacts must be continuously evaluated for policy violations. Although recent advances in large…

Many sets of ethics principles for responsible AI have been proposed to allay concerns about misuse and abuse of AI/ML systems. The underlying aspects of such sets of principles include privacy, accuracy, fairness, robustness,…

计算机与社会 · 计算机科学 2024-09-09 Conrad Sanderson , David Douglas , Qinghua Lu

In recent years, social media companies have grappled with defining and enforcing content moderation policies surrounding political content on their platforms, due in part to concerns about political bias, disinformation, and polarization.…

人机交互 · 计算机科学 2023-05-25 Jacob Thebault-Spieker , Sukrit Venkatagiri , Naomi Mine , Kurt Luther

Society is showing signs of strong ideological polarization. When pushed to seek perspectives different from their own, people often reject diverse ideas or find them unfathomable. Work has shown that framing controversial issues using the…

计算机与社会 · 计算机科学 2022-01-24 Jessica Wang , Amy Zhang , David Karger

Social platforms have revolutionized information sharing, but also accelerated the dissemination of harmful and policy-violating content. To ensure safety and compliance at scale, moderation systems must go beyond efficiency and offer…

计算与语言 · 计算机科学 2026-01-09 Anqi Li , Wenwei Jin , Jintao Tong , Pengda Qin , Weijia Li , Guo Lu

Moderation of reader comments is a significant problem for online news platforms. Here, we experiment with models for automatic moderation, using a dataset of comments from a popular Croatian newspaper. Our analysis shows that while…

计算与语言 · 计算机科学 2021-09-22 Elaine Zosa , Ravi Shekhar , Mladen Karan , Matthew Purver

Moral alignment has emerged as a widely adopted approach for regulating the behavior of pretrained language models (PLMs), typically through fine-tuning on curated datasets. Gender stereotype mitigation is a representational task within the…

计算与语言 · 计算机科学 2025-11-21 Guangliang Liu , Bocheng Chen , Han Zi , Xitong Zhang , Kristen Marie Johnson