中文
相关论文

相关论文: On Hate Scaling Laws For Data-Swamps

200 篇论文

Well-annotated data is a prerequisite for good Natural Language Processing models. Too often, though, annotation decisions are governed by optimizing time or annotator agreement. We make a case for nuanced efforts in an interdisciplinary…

计算与语言 · 计算机科学 2022-10-31 Federico Bianchi , Stefanie Anja Hills , Patricia Rossini , Dirk Hovy , Rebekah Tromble , Nava Tintarev

Large Language Models (LLM) have made significant advances in the recent past becoming more mainstream in Artificial Intelligence (AI) enabled human-facing applications. However, LLMs often generate stereotypical output inherited from…

计算与语言 · 计算机科学 2023-11-27 Wu Zekun , Sahan Bulathwela , Adriano Soares Koshiyama

With the exponential rise in user-generated web content on social media, the proliferation of abusive languages towards an individual or a group across the different sections of the internet is also rapidly increasing. It is very…

计算与语言 · 计算机科学 2021-03-24 Prashant Kapil , Asif Ekbal

Commercial Large Language Models (LLMs) have recently incorporated memory features to deliver personalised responses. This memory retains details such as user demographics and individual characteristics, allowing LLMs to adjust their…

计算与语言 · 计算机科学 2025-05-06 Paloma Piot , Patricia Martín-Rodilla , Javier Parapar

Large language models (LLMs) offer promising opportunities for organizational research. However, their built-in moderation systems can create problems when researchers try to analyze harmful content, often refusing to follow certain…

人工智能 · 计算机科学 2025-06-23 Mustafa Akben , Aaron Satko

Large language models (LLMs) are increasingly deployed in high-stakes domains, where rare but severe failures can result in irreversible harm. However, prevailing evaluation benchmarks often reduce complex social risk to mean-centered…

计算与语言 · 计算机科学 2026-01-30 Alok Abhishek , Tushar Bandopadhyay , Lisa Erickson

This paper investigates how hate speech varies in systematic ways according to the identities it targets. Across multiple hate speech datasets annotated for targeted identities, we find that classifiers trained on hate speech targeting…

计算与语言 · 计算机科学 2022-12-08 Michael Miller Yoder , Lynnette Hui Xian Ng , David West Brown , Kathleen M. Carley

With the freedom of communication provided in online social media, hate speech has increasingly generated. This leads to cyber conflicts affecting social life at the individual and national levels. As a result, hateful content…

计算与语言 · 计算机科学 2022-09-16 Khouloud Mnassri , Praboda Rajapaksha , Reza Farahbakhsh , Noel Crespi

Preference learning is a central component for aligning current LLMs, but this process can be vulnerable to data poisoning attacks. To address this concern, we introduce PoisonBench, a benchmark for evaluating large language models'…

密码学与安全 · 计算机科学 2025-06-09 Tingchen Fu , Mrinank Sharma , Philip Torr , Shay B. Cohen , David Krueger , Fazl Barez

Hate speech has grown significantly on social media, causing serious consequences for victims of all demographics. Despite much attention being paid to characterize and detect discriminatory speech, most work has focused on explicit or…

Machine learning models are widely adopted in scenarios that directly affect people. The development of software systems based on these models raises societal and legal concerns, as their decisions may lead to the unfair treatment of…

机器学习 · 计算机科学 2019-10-08 Inês Valentim , Nuno Lourenço , Nuno Antunes

Online harassment in the form of hate speech has been on the rise in recent years. Addressing the issue requires a combination of content moderation by people, aided by automatic detection methods. As content moderation is itself harmful to…

计算与语言 · 计算机科学 2021-08-03 Sheikh Muhammad Sarwar , Vanessa Murdock

Geographical, gender and stereotypical biases in computer vision models pose significant challenges to their performance and fairness. {In this study, we present an approach named FaceSaliencyAug aimed at addressing the gender bias in}…

计算机视觉与模式识别 · 计算机科学 2024-10-21 Teerath Kumar , Alessandra Mileo , Malika Bendechache

Following on recent advances in large language models (LLMs) and subsequent chat models, a new wave of large vision-language models (LVLMs) has emerged. Such models can incorporate images as input in addition to text, and perform tasks such…

计算机与社会 · 计算机科学 2024-02-09 Kathleen C. Fraser , Svetlana Kiritchenko

Amidst the rapid normalization of generative artificial intelligence (GAI), intelligent systems have come to dominate political discourse across information media. However, internalized political biases stemming from training data skews,…

计算与语言 · 计算机科学 2025-11-04 Nathan Junzi Chen

We investigate the potential for nationality biases in natural language processing (NLP) models using human evaluation methods. Biased NLP models can perpetuate stereotypes and lead to algorithmic discrimination, posing a significant…

In the past few years, there has been a significant rise in toxic and hateful content on various social media platforms. Recently Black Lives Matter movement came into the picture, causing an avalanche of user generated responses on the…

计算与语言 · 计算机科学 2021-08-31 Sumit Kumar , Raj Ratn Pranesh

Many existing works have made great strides towards reducing racial bias in face recognition. However, most of these methods attempt to rectify bias that manifests in models during training instead of directly addressing a major source of…

计算机视觉与模式识别 · 计算机科学 2021-10-06 Matthew Gwilliam , Srinidhi Hegde , Lade Tinubu , Alex Hanson

With the rise of online hate speech, automatic detection of Hate Speech, Offensive texts as a natural language processing task is getting popular. However, very little research has been done to detect unintended social bias from these toxic…

计算与语言 · 计算机科学 2022-10-24 Nihar Sahoo , Himanshu Gupta , Pushpak Bhattacharyya

Hate speech detection is a crucial area of research in natural language processing, essential for ensuring online community safety. However, detecting implicit hate speech, where harmful intent is conveyed in subtle or indirect ways,…

计算与语言 · 计算机科学 2025-04-17 Yumin Kim , Hwanhee Lee