中文
相关论文

相关论文: Watching the Watchers: A Comparative Fairness Audi…

200 篇论文

Commercial content moderation APIs are marketed as scalable solutions to combat online hate speech. However, the reliance on these APIs risks both silencing legitimate speech, called over-moderation, and failing to protect online platforms…

人机交互 · 计算机科学 2025-03-04 David Hartmann , Amin Oueslati , Dimitri Staufer , Lena Pohlmann , Simon Munzert , Hendrik Heuer

Online harms are a growing problem in digital spaces, putting user safety at risk and reducing trust in social media platforms. One of the most persistent forms of harm is hate speech. To address this, we need tools that combine the speed…

计算与语言 · 计算机科学 2025-09-03 Paloma Piot , Diego Sánchez , Javier Parapar

To meet the demands of content moderation, online platforms have resorted to automated systems. Newer forms of real-time engagement($\textit{e.g.}$, users commenting on live streams) on platforms like Twitch exert additional pressures on…

计算与语言 · 计算机科学 2025-06-11 Prarabdh Shukla , Wei Yin Chong , Yash Patel , Brennan Schaffner , Danish Pruthi , Arjun Bhagoji

Content moderation faces a challenging task as social media's ability to spread hate speech contrasts with its role in promoting global connectivity. With rapidly evolving slang and hate speech, the adaptability of conventional deep…

机器学习 · 计算机科学 2024-04-18 Paras Sheth , Tharindu Kumarage , Raha Moraffah , Aman Chadha , Huan Liu

Automatic content moderation is crucial to ensuring safety in social media. Language Model-based classifiers are being increasingly adopted for this task, but it has been shown that they perpetuate racial and social biases. Even if several…

计算与语言 · 计算机科学 2026-03-12 Alessandra Urbinati , Mirko Lai , Simona Frenda , Marco Antonio Stranisci

Social media platforms curate access to information and opportunities, and so play a critical role in shaping public discourse today. The opaque nature of the algorithms these platforms use to curate content raises societal questions. Prior…

计算机与社会 · 计算机科学 2023-03-08 Basileal Imana , Aleksandra Korolova , John Heidemann

To protect users from massive hateful content, existing works studied automated hate speech detection. Despite the existing efforts, one question remains: do automated hate speech detectors conform to social media content policies? A…

AI Safety Moderation (ASM) classifiers are designed to moderate content on social media platforms and to serve as guardrails that prevent Large Language Models (LLMs) from being fine-tuned on unsafe inputs. Owing to their potential for…

计算与语言 · 计算机科学 2025-01-24 Akshit Achara , Anshuman Chhabra

Machine learning (ML)-based content moderation tools are essential to keep online spaces free from hateful communication. Yet, ML tools can only be as capable as the quality of the data they are trained on allows them. While there is…

计算与语言 · 计算机科学 2024-06-14 Zehui Yu , Indira Sen , Dennis Assenmacher , Mattia Samory , Leon Fröhling , Christina Dahn , Debora Nozza , Claudia Wagner

Hate speech remains a persistent and unresolved challenge in online platforms. Content moderators, working on the front lines to review user-generated content and shield viewers from hate speech, often find themselves unprotected from the…

人机交互 · 计算机科学 2025-08-04 Subin Park , Jeonghyun Kim , Jeanne Choi , Joseph Seering , Uichin Lee , Sung-Ju Lee

Cyberbullying, which often has a deeply negative impact on the victim, has grown as a serious issue in Online Social Networks. Recently, researchers have created automated machine learning algorithms to detect Cyberbullying using social and…

社会与信息网络 · 计算机科学 2019-05-10 Vivek Singh , Connor Hofenbitzer

The proliferation of harmful content on online platforms is a major societal problem, which comes in many different forms including hate speech, offensive language, bullying and harassment, misinformation, spam, violence, graphic content,…

Unbiased data collection is essential to guaranteeing fairness in artificial intelligence models. Implicit bias, a form of behavioral conditioning that leads us to attribute predetermined characteristics to members of certain groups and…

人工智能 · 计算机科学 2020-03-03 Rupam Acharyya , Shouman Das , Ankani Chattoraj , Oishani Sengupta , Md Iftekar Tanveer

In this work, we introduce a novel metric for auditing group fairness in ranked lists. Our approach offers two benefits compared to the state of the art. First, we offer a blueprint for modeling of user attention. Rather than assuming a…

计算机与社会 · 计算机科学 2019-05-14 Piotr Sapiezynski , Wesley Zeng , Ronald E. Robertson , Alan Mislove , Christo Wilson

Social media platforms provide users the freedom of expression and a medium to exchange information and express diverse opinions. Unfortunately, this has also resulted in the growth of abusive content with the purpose of discriminating…

计算与语言 · 计算机科学 2021-07-01 Sohail Akhtar , Valerio Basile , Viviana Patti

Automated content moderation has long been used to help identify and filter undesired user-generated content online. But such systems have a history of incorrectly flagging content by and about marginalized identities for removal.…

The rise of hate speech on online platforms has led to an urgent need for effective content moderation. However, the subjective and multi-faceted nature of hateful online content, including implicit hate speech, poses significant challenges…

计算与语言 · 计算机科学 2023-03-17 Uma Gunturi , Xiaohan Ding , Eugenia H. Rho

Hateful speech detection is a key component of content moderation, yet current evaluation frameworks rarely assess why a text is deemed hateful. We introduce \textsf{HateXScore}, a four-component metric suite designed to evaluate the…

计算与语言 · 计算机科学 2026-01-21 Yujia Hu , Roy Ka-Wei Lee

Online hate speech is associated with substantial social harms, yet it remains unclear how consistently platforms enforce hate speech policies or whether enforcement is feasible at scale. We address these questions through a global audit of…

Online hate is a growing concern on many social media platforms and other sites. To combat it, technology companies are increasingly identifying and sanctioning `hateful users' rather than simply moderating hateful content. Yet, most…

社会与信息网络 · 计算机科学 2021-03-23 Zo Ahmed , Bertie Vidgen , Scott A. Hale
‹ 上一页 1 2 3 10 下一页 ›