中文
相关论文

相关论文: Moderating Harm: Benchmarking Large Language Model…

200 篇论文

The exponential growth of social media has profoundly transformed how information is created, disseminated, and absorbed, exceeding any precedent in the digital age. Regrettably, this explosion has also spawned a significant increase in the…

计算与语言 · 计算机科学 2025-03-03 Hongzhan Lin , Ziyang Luo , Bo Wang , Ruichao Yang , Jing Ma

Online marketplaces, while revolutionizing global commerce, have inadvertently facilitated the proliferation of illicit activities, including drug trafficking, counterfeit sales, and cybercrimes. Traditional content moderation methods such…

计算与语言 · 计算机科学 2026-03-06 Quoc Khoa Tran , Thanh Thi Nguyen , Campbell Wilson

Programmers are turning to AI coding assistants to answer questions about their code. Benchmarks are needed to soundly evaluate these systems and understand their performance. To enable such a study, we curate a benchmark of real-world…

软件工程 · 计算机科学 2026-05-06 Ferida Mohammed , Fatma Ayad , Petros Maniatis , Satish Chandra , Elizabeth Dinella

This research offers a unique evaluation of how AI systems interpret the digital language of Generation Alpha (Gen Alpha, born 2010-2024). As the first cohort raised alongside AI, Gen Alpha faces new forms of online risk due to immersive…

计算机与社会 · 计算机科学 2025-05-19 Manisha Mehta , Fausto Giunchiglia

Social media has become an essential part of the daily routines of children and adolescents. Moreover, enormous efforts have been made to ensure the psychological and emotional well-being of young users as well as their safety when…

计算机与社会 · 计算机科学 2021-03-17 Sultan Alshamrani , Ahmed Abusnaina , Mohammed Abuhamad , Daehun Nyang , David Mohaisen

This paper explores a novel method for enhancing binary classification models that assess code comment quality, leveraging Generative Artificial Intelligence to elevate model performance. By integrating 1,437 newly generated code-comment…

软件工程 · 计算机科学 2024-10-30 Seetharam Killivalavan , Durairaj Thenmozhi

Social media platforms utilize Machine Learning (ML) and Artificial Intelligence (AI) powered recommendation algorithms to maximize user engagement, which can result in inadvertent exposure to harmful content. Current moderation efforts,…

计算与语言 · 计算机科学 2025-05-30 Rajvardhan Oak , Muhammad Haroon , Claire Jo , Magdalena Wojcieszak , Anshuman Chhabra

Detecting hateful content is a challenging and important problem. Automated tools, like machine-learning models, can help, but they require continuous training to adapt to the ever-changing landscape of social media. In this work, we…

计算与语言 · 计算机科学 2025-11-06 Jay Patel , Hrudayangam Mehta , Jeremy Blackburn

This paper presents an eight-week observational comparison of 68 single-file HTML generations collected across 17 public experiments in the "HTML AI Battle" project between December 10, 2025 and February 4, 2026. Four reasoning model…

软件工程 · 计算机科学 2026-05-11 Diego Cabezas Palacios

This study presents the first large-scale, side-by-side comparison of contemporary Large Language Models (LLMs) in the automated grading of programming assignments. Drawing on over 6,000 student submissions collected across four years of an…

计算机与社会 · 计算机科学 2025-10-01 Marcin Jukiewicz

Social media are shifting towards pluralism -- community-governed platforms where groups define their own norms. What violates rules in one community may be perfectly acceptable in another. Can AI models help moderate such pluralistic…

计算与语言 · 计算机科学 2026-05-19 Zoher Kachwala , Bao Tran Truong , Rasika Muralidharan , Haewoon Kwak , Jisun An , Filippo Menczer

On social media platforms, hateful and offensive language negatively impact the mental well-being of users and the participation of people from diverse backgrounds. Automatic methods to detect offensive language have largely relied on…

计算与语言 · 计算机科学 2022-01-26 Rishav Hada , Sohi Sudhir , Pushkar Mishra , Helen Yannakoudakis , Saif M. Mohammad , Ekaterina Shutova

Cyberbullying is a pervasive problem in online communities. To identify cyberbullying cases in large-scale social networks, content moderators depend on machine learning classifiers for automatic cyberbullying detection. However, existing…

社会与信息网络 · 计算机科学 2020-04-07 Caleb Ziems , Ymir Vigfusson , Fred Morstatter

The problems of online hate speech and cyberbullying have significantly worsened since the increase in popularity of social media platforms such as YouTube and Twitter (X). Natural Language Processing (NLP) techniques have proven to provide…

计算与语言 · 计算机科学 2024-03-18 Sargam Yadav , Abhishek Kaushik , Kevin McDaid

Short-video platforms have become major channels for misinformation, where deceptive claims frequently leverage visual experiments and social cues. While Multimodal Large Language Models (MLLMs) have demonstrated impressive reasoning…

计算与语言 · 计算机科学 2026-05-19 Jen-tse Huang , Chang Chen , Shiyang Lai , Wenxuan Wang , Michelle R. Kaufman , Mark Dredze

Artificial Intelligence (AI)-generated feedback in educational settings has garnered considerable attention due to its potential to enhance learning outcomes. However, a comprehensive understanding of the linguistic characteristics of…

计算与语言 · 计算机科学 2025-05-01 Antoun Yaacoub , Zainab Assaghir , Lionel Prevost , Jérôme Da-Rugna

Online hate speech is associated with substantial social harms, yet it remains unclear how consistently platforms enforce hate speech policies or whether enforcement is feasible at scale. We address these questions through a global audit of…

Toxic language, such as hate speech, can deter users from participating in online communities and enjoying popular platforms. Previous approaches to detecting toxic language and norm violations have been primarily concerned with…

计算与语言 · 计算机科学 2023-10-10 Jihyung Moon , Dong-Ho Lee , Hyundong Cho , Woojeong Jin , Chan Young Park , Minwoo Kim , Jonathan May , Jay Pujara , Sungjoon Park

This paper tries to address the problem of abusive comment detection in low-resource indic languages. Abusive comments are statements that are offensive to a person or a group of people. These comments are targeted toward individuals…

计算与语言 · 计算机科学 2022-04-22 Shantanu Patankar , Omkar Gokhale , Onkar Litake , Aditya Mandke , Dipali Kadam

Content moderation is a central mechanism through which platforms attempt to balance user engagement with community governance. Yet existing research has largely treated moderation as a uniform intervention, overlooking how moderator…

计算机与社会 · 计算机科学 2026-05-18 Siyi Zhou , Lindsay Young , Marlon Twyman , Emilio Ferrara