English
Related papers

Related papers: Reliable Decision from Multiple Subtasks through T…

200 papers

The sheer volume of online user-generated content has rendered content moderation technologies essential in order to protect digital platform audiences from content that may cause anxiety, worry, or concern. Despite the efforts towards…

Computer Vision and Pattern Recognition · Computer Science 2022-12-02 Ioannis Sarridis , Christos Koutlis , Olga Papadopoulou , Symeon Papadopoulos

Toxicity annotators and content moderators often default to mental shortcuts when making decisions. This can lead to subtle toxicity being missed, and seemingly toxic but harmless content being over-detected. We introduce BiasX, a framework…

Computation and Language · Computer Science 2023-05-24 Yiming Zhang , Sravani Nanduri , Liwei Jiang , Tongshuang Wu , Maarten Sap

We present Moderator, a policy-based model management system that allows administrators to specify fine-grained content moderation policies and modify the weights of a text-to-image (TTI) model to make it significantly more challenging for…

Cryptography and Security · Computer Science 2024-09-13 Peiran Wang , Qiyu Li , Longxuan Yu , Ziyao Wang , Ang Li , Haojian Jin

This paper analyzes the community safety guidelines of five text-to-image (T2I) generation platforms and audits five T2I models, focusing on prompts related to the representation of humans in areas that might lead to societal stigma. While…

Computers and Society · Computer Science 2024-09-27 Piera Riccio , Georgina Curto , Nuria Oliver

Platform content moderation applies explicit policy rules and context-dependent conditions to decide whether user content is allowed, restricted, or removed. A correct moderation outcome must therefore depend on which rules a case…

Artificial Intelligence · Computer Science 2026-05-11 Zhifeng Lu , Dianyuan Wang , Yuhu Shang , Zhenbo Xu

Accurately estimating how users respond to moderation interventions is paramount for developing effective and user-centred moderation strategies. However, this requires a clear understanding of which user characteristics are associated with…

Computers and Society · Computer Science 2025-10-24 Benedetta Tessa , Alejandro Moreo , Stefano Cresci , Tiziano Fagni , Fabrizio Sebastiani

The increasing scale and complexity of online platforms raises critical policy questions around harmful content, digital well-being, and user autonomy. Traditional content moderation systems rely on centralised, top-down rules, often…

Computers and Society · Computer Science 2026-05-05 Ewelina Gajewska , Michal Wawer , Katarzyna Budzynska , Jaroslaw A. Chudziak

User-generated content (UGC) on social media platforms is vulnerable to incitements and manipulations, necessitating effective regulations. To address these challenges, those platforms often deploy automated content moderators tasked with…

Machine Learning · Computer Science 2025-07-29 Saba Ahmadi , Avrim Blum , Haifeng Xu , Fan Yao

Social media platforms have diverse content moderation policies, with many prominent actors hesitant to impose strict regulations. A key reason for this reluctance could be the competitive advantage that comes with lax regulation. A popular…

Computer Science and Game Theory · Computer Science 2024-02-17 So Sasaki , Cédric Langbort

With the growth of social media and large language models, content moderation has become crucial. Many existing datasets lack adequate representation of different groups, resulting in unreliable assessments. To tackle this, we propose a…

Computation and Language · Computer Science 2024-12-19 Shanu Kumar , Gauri Kholkar , Saish Mendke , Anubhav Sadana , Parag Agrawal , Sandipan Dandapat

The proliferation of harmful content on online platforms is a major societal problem, which comes in many different forms including hate speech, offensive language, bullying and harassment, misinformation, spam, violence, graphic content,…

The growth of online platforms and user content requires strong content moderation systems that can handle complex inputs from various media types. While large language models (LLMs) are effective, their high computational cost and latency…

Computation and Language · Computer Science 2026-04-09 Shutong Zhang , Dylan Zhou , Yinxiao Liu , Yang Yang , Huiwen Luo , Wenfei Zou

Content moderation on a global scale must navigate a complex array of local cultural distinctions, which can hinder effective enforcement. While global policies aim for consistency and broad applicability, they often miss the subtleties of…

Toxicity detection algorithms, originally designed with reactive content moderation in mind, are increasingly being deployed into proactive end-user interventions to moderate content. Through a socio-technical lens and focusing on contexts…

Human-Computer Interaction · Computer Science 2025-02-25 Mark Warner , Angelika Strohmayer , Matthew Higgs , Lynne Coventry

With the rapid rise of short-form videos, TikTok has become one of the most influential platforms among children and teenagers, but also a source of harmful content that can affect their perception and behavior. Such content, often subtle…

Computation and Language · Computer Science 2025-11-25 Dat Thanh Nguyen , Nguyen Hung Lam , Anh Hoang-Thi Nguyen , Trong-Hop Do

There is an ongoing debate about how to moderate toxic speech on social media and the impact of content moderation on online discourse. This paper proposes and validates a methodology for measuring the content-moderation-induced distortions…

Social and Information Networks · Computer Science 2026-03-04 Mahyar Habibi , Dirk Hovy , Carlo Schwarz

Most social media users come from the Global South, where harmful content usually appears in local languages. Yet, AI-driven moderation systems struggle with low-resource languages spoken in these regions. Through semi-structured interviews…

Computation and Language · Computer Science 2025-08-06 Farhana Shahid , Mona Elswah , Aditya Vashistha

The exponential growth of social media platforms such as Twitter and Facebook has revolutionized textual communication and textual content publication in human society. However, they have been increasingly exploited to propagate toxic…

Computation and Language · Computer Science 2023-02-14 Wenxuan Wang , Jen-tse Huang , Weibin Wu , Jianping Zhang , Yizhan Huang , Shuqing Li , Pinjia He , Michael Lyu

Facebook and Twitter recently announced community-based review platforms to address misinformation. We provide an overview of the potential affordances of such community-based approaches to content moderation based on past research and…

Social and Information Networks · Computer Science 2023-01-05 Taha Yasseri , Filippo Menczer

Social media are shifting towards pluralism -- community-governed platforms where groups define their own norms. What violates rules in one community may be perfectly acceptable in another. Can AI models help moderate such pluralistic…

Computation and Language · Computer Science 2026-05-19 Zoher Kachwala , Bao Tran Truong , Rasika Muralidharan , Haewoon Kwak , Jisun An , Filippo Menczer