English
Related papers

Related papers: GMP: A Benchmark for Content Moderation under Co-o…

200 papers

Evaluating the value alignment of large language models (LLMs) has traditionally relied on single-sentence adversarial prompts, which directly probe models with ethically sensitive or controversial questions. However, with the rapid…

Computation and Language · Computer Science 2025-03-31 Yazhou Zhang , Qimeng Liu , Qiuchi Li , Peng Zhang , Jing Qin

We study the impact of content moderation policies in online communities. In our theoretical model, a platform chooses a content moderation policy and individuals choose whether or not to participate in the community according to the…

Data Structures and Algorithms · Computer Science 2023-10-17 Cynthia Dwork , Chris Hays , Jon Kleinberg , Manish Raghavan

Online abuse is becoming an increasingly prevalent issue in modern-day society, with 41 percent of Americans having experienced online harassment in some capacity in 2021. People who identify as women, in particular, can be subjected to a…

Human-Computer Interaction · Computer Science 2023-01-19 Sarah Barrington

Extensive efforts in automated approaches for content moderation have been focused on developing models to identify toxic, offensive, and hateful content with the aim of lightening the load for moderators. Yet, it remains uncertain whether…

Computation and Language · Computer Science 2024-11-14 Yang Trista Cao , Lovely-Frances Domingo , Sarah Ann Gilbert , Michelle Mazurek , Katie Shilton , Hal Daumé

Benchmarks are a cornerstone of modern machine learning, enabling reproducibility, comparison, and scientific progress. However, AI benchmarks are increasingly complex, requiring dynamic, AI-focused workflows. Rapid evolution in model…

As large language models (LLMs) evolve into autonomous agents capable of acting in open-ended environments, ensuring behavioral alignment with human values becomes a critical safety concern. Existing benchmarks, focused on static,…

Computation and Language · Computer Science 2026-03-10 Weixiang Zhao , Haozhen Li , Yanyan Zhao , xuda zhi , Yongbo Huang , Hao He , Bing Qin , Ting Liu

Online platforms and communities establish their own norms that govern what behavior is acceptable within the community. Substantial effort in NLP has focused on identifying unacceptable behaviors and, recently, on forecasting them before…

Computation and Language · Computer Science 2021-10-12 Chan Young Park , Julia Mendelsohn , Karthik Radhakrishnan , Kinjal Jain , Tushar Kanakagiri , David Jurgens , Yulia Tsvetkov

Artificial intelligence (AI) model creators commonly attach restrictive terms of use to both their models and their outputs. These terms typically prohibit activities ranging from creating competing AI models to spreading disinformation.…

Computers and Society · Computer Science 2024-12-11 Peter Henderson , Mark A. Lemley

Content moderation systems are typically evaluated by measuring agreement with human labels. In rule-governed environments this assumption fails: multiple decisions may be logically consistent with the governing policy, and agreement…

Artificial Intelligence · Computer Science 2026-04-24 Michael O'Herlihy , Rosa Català

Social media platforms have implemented automated content moderation tools to preserve community norms and mitigate online hate and harassment. Recently, these platforms have started to offer Personalized Content Moderation (PCM), granting…

Social and Information Networks · Computer Science 2024-05-17 Necdet Gurkan , Mohammed Almarzouq , Pon Rahul Murugaraj

Large language models are increasingly deployed not as single assistants but as committees whose members deliberate and then vote or synthesize a decision. Such systems are often expected to be more robust than individual models. We show…

Artificial Intelligence · Computer Science 2026-04-07 Hajime Shimao , Warut Khern-am-nuai , Sung Joo Kim

Online platforms are seeing increasing amounts of AI-generated content -- text and other forms of media that are made or co-created with generative AI. This trend suggests platforms may need to establish governance frameworks, including…

Human-Computer Interaction · Computer Science 2026-03-16 Lan Gao , Abani Ahmed , Oscar Chen , Margaux Reyl , Zayna Cheema , Nick Feamster , Chenhao Tan , Kurt Thomas , Marshini Chetty

Generative AI (GenAI) is increasingly being integrated into the online ecosystem, including online health communities (OHCs), where people with diverse health conditions exchange social support. For example, in OHCs, support providers are…

Benchmarks play a significant role in how technology companies communicate about model capabilities and how researchers and the public understand generative AI systems. However, existing benchmarks have been criticized for their failure to…

Human-Computer Interaction · Computer Science 2026-04-29 Charlotte Li , Nick Hagar , Sachita Nishal , Jeremy Gilbert , Nick Diakopoulos

This Article examines the constitutional status of AI-mediated communication under the First Amendment. Social media platforms, increasingly integrated with generative AI systems, now function as core public communication infrastructures.…

Computers and Society · Computer Science 2026-03-03 Yiyang Mei

Content moderation remains a critical yet challenging task for large-scale user-generated video platforms, especially in livestreaming environments where moderation must be timely, multimodal, and robust to evolving forms of unwanted…

Computer Vision and Pattern Recognition · Computer Science 2026-01-29 Wei Chee Yew , Hailun Xu , Sanjay Saha , Xiaotian Fan , Hiok Hian Ong , David Yuchen Wang , Kanchan Sarkar , Zhenheng Yang , Danhui Guan

AI-generated counterspeech offers a promising and scalable strategy to curb online toxicity through direct replies that promote civil discourse. However, current counterspeech is one-size-fits-all, lacking adaptation to the moderation…

Human-Computer Interaction · Computer Science 2025-02-10 Lorenzo Cima , Alessio Miaschi , Amaury Trujillo , Marco Avvenuti , Felice Dell'Orletta , Stefano Cresci

This study establishes a novel framework for systematically evaluating the moral reasoning capabilities of large language models (LLMs) as they increasingly integrate into critical societal domains. Current assessment methodologies lack the…

Computers and Society · Computer Science 2025-05-05 Junfeng Jiao , Saleh Afroogh , Abhejay Murali , Kevin Chen , David Atkinson , Amit Dhurandhar

As autonomous AI agents are increasingly deployed in high-stakes environments, ensuring their safety and alignment with human values is becoming a practical deployment concern. Current benchmarks for AI agents primarily evaluate refusal of…

Artificial Intelligence · Computer Science 2026-05-12 Miles Q. Li , Benjamin C. M. Fung , Martin Weiss , Pulei Xiong , Khalil Al-Hussaeni , Claude Fachkha

Generative artificial intelligence (AI) is increasingly integrated into the online platforms where humans exchange opinions; large language models (LLMs) now polish users' posts on LinkedIn and provide context for content shared on X. While…

Computers and Society · Computer Science 2026-05-18 Stratis Tsirtsis , Kai Rawal , Chris Russell , Brent Mittelstadt , Sandra Wachter