中文
相关论文

相关论文: GMP: A Benchmark for Content Moderation under Co-o…

200 篇论文

Evaluating the value alignment of large language models (LLMs) has traditionally relied on single-sentence adversarial prompts, which directly probe models with ethically sensitive or controversial questions. However, with the rapid…

计算与语言 · 计算机科学 2025-03-31 Yazhou Zhang , Qimeng Liu , Qiuchi Li , Peng Zhang , Jing Qin

We study the impact of content moderation policies in online communities. In our theoretical model, a platform chooses a content moderation policy and individuals choose whether or not to participate in the community according to the…

数据结构与算法 · 计算机科学 2023-10-17 Cynthia Dwork , Chris Hays , Jon Kleinberg , Manish Raghavan

Online abuse is becoming an increasingly prevalent issue in modern-day society, with 41 percent of Americans having experienced online harassment in some capacity in 2021. People who identify as women, in particular, can be subjected to a…

人机交互 · 计算机科学 2023-01-19 Sarah Barrington

Extensive efforts in automated approaches for content moderation have been focused on developing models to identify toxic, offensive, and hateful content with the aim of lightening the load for moderators. Yet, it remains uncertain whether…

计算与语言 · 计算机科学 2024-11-14 Yang Trista Cao , Lovely-Frances Domingo , Sarah Ann Gilbert , Michelle Mazurek , Katie Shilton , Hal Daumé

Benchmarks are a cornerstone of modern machine learning, enabling reproducibility, comparison, and scientific progress. However, AI benchmarks are increasingly complex, requiring dynamic, AI-focused workflows. Rapid evolution in model…

As large language models (LLMs) evolve into autonomous agents capable of acting in open-ended environments, ensuring behavioral alignment with human values becomes a critical safety concern. Existing benchmarks, focused on static,…

计算与语言 · 计算机科学 2026-03-10 Weixiang Zhao , Haozhen Li , Yanyan Zhao , xuda zhi , Yongbo Huang , Hao He , Bing Qin , Ting Liu

Online platforms and communities establish their own norms that govern what behavior is acceptable within the community. Substantial effort in NLP has focused on identifying unacceptable behaviors and, recently, on forecasting them before…

Artificial intelligence (AI) model creators commonly attach restrictive terms of use to both their models and their outputs. These terms typically prohibit activities ranging from creating competing AI models to spreading disinformation.…

计算机与社会 · 计算机科学 2024-12-11 Peter Henderson , Mark A. Lemley

Content moderation systems are typically evaluated by measuring agreement with human labels. In rule-governed environments this assumption fails: multiple decisions may be logically consistent with the governing policy, and agreement…

人工智能 · 计算机科学 2026-04-24 Michael O'Herlihy , Rosa Català

Social media platforms have implemented automated content moderation tools to preserve community norms and mitigate online hate and harassment. Recently, these platforms have started to offer Personalized Content Moderation (PCM), granting…

社会与信息网络 · 计算机科学 2024-05-17 Necdet Gurkan , Mohammed Almarzouq , Pon Rahul Murugaraj

Large language models are increasingly deployed not as single assistants but as committees whose members deliberate and then vote or synthesize a decision. Such systems are often expected to be more robust than individual models. We show…

人工智能 · 计算机科学 2026-04-07 Hajime Shimao , Warut Khern-am-nuai , Sung Joo Kim

Online platforms are seeing increasing amounts of AI-generated content -- text and other forms of media that are made or co-created with generative AI. This trend suggests platforms may need to establish governance frameworks, including…

人机交互 · 计算机科学 2026-03-16 Lan Gao , Abani Ahmed , Oscar Chen , Margaux Reyl , Zayna Cheema , Nick Feamster , Chenhao Tan , Kurt Thomas , Marshini Chetty

Generative AI (GenAI) is increasingly being integrated into the online ecosystem, including online health communities (OHCs), where people with diverse health conditions exchange social support. For example, in OHCs, support providers are…

Benchmarks play a significant role in how technology companies communicate about model capabilities and how researchers and the public understand generative AI systems. However, existing benchmarks have been criticized for their failure to…

人机交互 · 计算机科学 2026-04-29 Charlotte Li , Nick Hagar , Sachita Nishal , Jeremy Gilbert , Nick Diakopoulos

This Article examines the constitutional status of AI-mediated communication under the First Amendment. Social media platforms, increasingly integrated with generative AI systems, now function as core public communication infrastructures.…

计算机与社会 · 计算机科学 2026-03-03 Yiyang Mei

Content moderation remains a critical yet challenging task for large-scale user-generated video platforms, especially in livestreaming environments where moderation must be timely, multimodal, and robust to evolving forms of unwanted…

计算机视觉与模式识别 · 计算机科学 2026-01-29 Wei Chee Yew , Hailun Xu , Sanjay Saha , Xiaotian Fan , Hiok Hian Ong , David Yuchen Wang , Kanchan Sarkar , Zhenheng Yang , Danhui Guan

AI-generated counterspeech offers a promising and scalable strategy to curb online toxicity through direct replies that promote civil discourse. However, current counterspeech is one-size-fits-all, lacking adaptation to the moderation…

人机交互 · 计算机科学 2025-02-10 Lorenzo Cima , Alessio Miaschi , Amaury Trujillo , Marco Avvenuti , Felice Dell'Orletta , Stefano Cresci

This study establishes a novel framework for systematically evaluating the moral reasoning capabilities of large language models (LLMs) as they increasingly integrate into critical societal domains. Current assessment methodologies lack the…

计算机与社会 · 计算机科学 2025-05-05 Junfeng Jiao , Saleh Afroogh , Abhejay Murali , Kevin Chen , David Atkinson , Amit Dhurandhar

As autonomous AI agents are increasingly deployed in high-stakes environments, ensuring their safety and alignment with human values is becoming a practical deployment concern. Current benchmarks for AI agents primarily evaluate refusal of…

人工智能 · 计算机科学 2026-05-12 Miles Q. Li , Benjamin C. M. Fung , Martin Weiss , Pulei Xiong , Khalil Al-Hussaeni , Claude Fachkha

Generative artificial intelligence (AI) is increasingly integrated into the online platforms where humans exchange opinions; large language models (LLMs) now polish users' posts on LinkedIn and provide context for content shared on X. While…

计算机与社会 · 计算机科学 2026-05-18 Stratis Tsirtsis , Kai Rawal , Chris Russell , Brent Mittelstadt , Sandra Wachter