中文
相关论文

相关论文: Like trainer, like bot? Inheritance of bias in alg…

200 篇论文

When LLM-based multi-agent systems disagree, current practice treats this as noise to be resolved through consensus. We propose it can be signal. We focus on hate speech moderation, a domain where judgments depend on cultural context and…

多智能体系统 · 计算机科学 2026-04-07 Michał Wawer , Jarosław A. Chudziak

Collecting annotations from human raters often results in a trade-off between the quantity of labels one wishes to gather and the quality of these labels. As such, it is often only possible to gather a small amount of high-quality labels.…

机器学习 · 计算机科学 2021-10-05 Neel Nanda , Jonathan Uesato , Sven Gowal

The spread of online hate has become a significant problem for newspapers that host comment sections. As a result, there is growing interest in using machine learning and natural language processing for (semi-) automated abusive language…

计算与语言 · 计算机科学 2022-07-11 Lennart Justen , Kilian Müller , Marco Niemann , Jörg Becker

Large language models (LLMs) are increasingly used in content moderation systems, where ensuring fairness and neutrality is essential. In this study, we examine how persona adoption influences the consistency and fairness of harmful content…

计算与语言 · 计算机科学 2025-10-31 Stefano Civelli , Pietro Bernardelle , Nardiena A. Pratama , Gianluca Demartini

Artificial Intelligence has the capacity to amplify and perpetuate societal biases and presents profound ethical implications for society. Gender bias has been identified in the context of employment advertising and recruitment tools, due…

计算与语言 · 计算机科学 2020-05-19 Susan Leavy , Gerardine Meaney , Karen Wade , Derek Greene

With peak content moderation seemingly behind us, this paper revisits its punitive side. But instead of focusing on who is being (disproportionately) moderated, it focuses on the punishment itself and explores the question of how content…

计算机与社会 · 计算机科学 2026-05-14 Robert Grimm

As large language models (LLMs) are increasingly deployed in high-stakes settings, their ability to refuse ethically sensitive prompts-such as those involving hate speech or illegal activities-has become central to content moderation and…

人机交互 · 计算机科学 2025-05-22 Stefan Pasch

Online abuse is becoming an increasingly prevalent issue in modern-day society, with 41 percent of Americans having experienced online harassment in some capacity in 2021. People who identify as women, in particular, can be subjected to a…

人机交互 · 计算机科学 2023-01-19 Sarah Barrington

In the digital age, hate speech poses a threat to the functioning of social media platforms as spaces for public discourse. Top-down approaches to moderate hate speech encounter difficulties due to conflicts with freedom of expression and…

计算机与社会 · 计算机科学 2025-11-24 Jana Lasser , Alina Herderich , Joshua Garland , Segun Taofeek Aroyehun , David Garcia , Mirta Galesic

Online communities have gained considerable importance in recent years due to the increasing number of people connected to the Internet. Moderating user content in online communities is mainly performed manually, and reducing the workload…

信息检索 · 计算机科学 2019-01-16 Etienne Papegnies , Vincent Labatut , Richard Dufour , Georges Linares

Effective content moderation systems require explicit classification criteria, yet online communities like subreddits often operate with diverse, implicit standards. This work introduces a novel approach to identify and extract these…

计算与语言 · 计算机科学 2025-09-04 Youngwoo Kim , Himanshu Beniwal , Steven L. Johnson , Thomas Hartvigsen

Social bots play a significant role in many online social networks (OSN) as they imitate human behavior. This fact raises difficult questions about their capabilities and potential risks. Given the recent advances in Generative AI (GenAI),…

社会与信息网络 · 计算机科学 2024-05-06 Shaghayegh Najari , Davood Rafiee , Mostafa Salehi , Reza Farahbakhsh

Though majority vote among annotators is typically used for ground truth labels in natural language processing, annotator disagreement in tasks such as hate speech detection may reflect differences in opinion across groups, not noise. Thus,…

计算与语言 · 计算机科学 2024-03-19 Eve Fleisig , Rediet Abebe , Dan Klein

During deliberation processes, mediators and facilitators typically need to select a small and representative set of opinions later used to produce digestible reports for stakeholders. In online deliberation platforms, algorithmic selection…

计算机与社会 · 计算机科学 2026-02-18 Salim Hafid , Manon Berriche , Jean-Philippe Cointet

Today, social media platforms are significant sources of news and political communication, but their role in spreading misinformation has raised significant concerns. In response, these platforms have implemented various content moderation…

计算机与社会 · 计算机科学 2026-04-21 Saeedeh Mohammadi , Taha Yasseri

The flow of information reaching us via the online media platforms is optimized not by the information content or relevance but by popularity and proximity to the target. This is typically performed in order to maximise platform usage. As a…

物理与社会 · 物理学 2019-06-19 Alina Sîrbu , Dino Pedreschi , Fosca Giannotti , János Kertész

The increasing sophistication of large language models (LLMs) has sparked growing concerns regarding their potential role in exacerbating ideological polarization through the automated generation of persuasive and biased content. This study…

计算与语言 · 计算机科学 2025-06-18 . Pazzaglia , V. Vendetti , L. D. Comencini , F. Deriu , V. Modugno

Text classification is an important topic in the field of natural language processing. It has been preliminarily applied in information retrieval, digital library, automatic abstracting, text filtering, word semantic discrimination and many…

计算与语言 · 计算机科学 2023-12-20 Hao Li , Brandon Bennett

Modelling the complex dynamics of online social platforms is critical for addressing challenges such as hate speech and misinformation. While Discussion Transformers, which model conversations as graph structures, have emerged as a…

社会与信息网络 · 计算机科学 2026-02-04 Liam Hebert , Lucas Kopp , Robin Cohen

Platforms that support online commentary, from social networks to news sites, are increasingly leveraging machine learning to assist their moderation efforts. But this process does not typically provide feedback to the author that would…

计算与语言 · 计算机科学 2021-02-12 Leo Laugier , John Pavlopoulos , Jeffrey Sorensen , Lucas Dixon