中文
相关论文

相关论文: Mitigating Biases in Toxic Language Detection thro…

200 篇论文

Algorithmic bias often arises as a result of differential subgroup validity, in which predictive relationships vary across groups. For example, in toxic language detection, comments targeting different demographic groups can vary markedly…

机器学习 · 计算机科学 2023-03-08 Soumyajit Gupta , Sooyong Lee , Maria De-Arteaga , Matthew Lease

Due to the subtleness, implicity, and different possible interpretations perceived by different people, detecting undesirable content from text is a nuanced difficulty. It is a long-known risk that language models (LMs), once trained on…

计算与语言 · 计算机科学 2022-05-26 Yau-Shian Wang , Yingshan Chang

The rise of social networks has not only facilitated communication but also allowed the spread of harmful content. Although significant advances have been made in detecting toxic language in textual data, the exploration of concept-based…

计算与语言 · 计算机科学 2025-12-16 Samarth Garg , Divya Singh , Deeksha Varshney , Mamta

The rise of cyberbullying in social media platforms involving toxic comments has escalated the need for effective ways to monitor and moderate online interactions. Existing solutions of automated toxicity detection systems, are based on a…

人工智能 · 计算机科学 2026-02-11 Michail S. Alexiou , J. Sukarno Mertoguno

Text representation models are prone to exhibit a range of societal biases, reflecting the non-controlled and biased nature of the underlying pretraining data, which consequently leads to severe ethical issues and even bias amplification.…

计算与语言 · 计算机科学 2021-06-08 Soumya Barikeri , Anne Lauscher , Ivan Vulić , Goran Glavaš

With the ever-increasing cases of hate spread on social media platforms, it is critical to design abuse detection mechanisms to proactively avoid and control such incidents. While there exist methods for hate speech detection, they…

计算与语言 · 计算机科学 2020-01-17 Pinkesh Badjatiya , Manish Gupta , Vasudeva Varma

Warning: this paper contains content that maybe offensive or upsetting. Recent research in Natural Language Processing (NLP) has advanced the development of various toxicity detection models with the intention of identifying and mitigating…

计算与语言 · 计算机科学 2022-05-06 Ninareh Mehrabi , Ahmad Beirami , Fred Morstatter , Aram Galstyan

Now that AI-driven moderation has become pervasive in everyday life, we often hear claims that "the AI is biased". While this is often said jokingly, the light-hearted remark reflects a deeper concern. How can we be certain that an online…

计算与语言 · 计算机科学 2026-04-02 Subhojit Ghimire

Harmful content detection models tend to have higher false positive rates for content from marginalized groups. In the context of marginal abuse modeling on Twitter, such disproportionate penalization poses the risk of reduced visibility,…

计算与语言 · 计算机科学 2022-10-13 Kyra Yee , Alice Schoenauer Sebag , Olivia Redfield , Emily Sheng , Matthias Eck , Luca Belli

Natural language processing (NLP) models often replicate or amplify social bias from training data, raising concerns about fairness. At the same time, their black-box nature makes it difficult for users to recognize biased predictions and…

计算与语言 · 计算机科学 2026-02-12 Yifan Wang , Mayank Jobanputra , Ji-Ung Lee , Soyoung Oh , Isabel Valera , Vera Demberg

Modern toxic speech detectors are incompetent in recognizing disguised offensive language, such as adversarial attacks that deliberately avoid known toxic lexicons, or manifestations of implicit bias. Building a large annotated dataset for…

计算与语言 · 计算机科学 2020-10-08 Xiaochuang Han , Yulia Tsvetkov

Gender-bias stereotypes have recently raised significant ethical concerns in natural language processing. However, progress in detection and evaluation of gender bias in natural language understanding through inference is limited and…

计算与语言 · 计算机科学 2021-05-13 Shanya Sharma , Manan Dey , Koustuv Sinha

Toxicity detection algorithms, originally designed with reactive content moderation in mind, are increasingly being deployed into proactive end-user interventions to moderate content. Through a socio-technical lens and focusing on contexts…

人机交互 · 计算机科学 2025-02-25 Mark Warner , Angelika Strohmayer , Matthew Higgs , Lynne Coventry

Toxicity classification for voice heavily relies on the semantic content of speech. We propose a novel framework that utilizes cross-modal learning to integrate the semantic embedding of text into a multilabel speech toxicity classifier…

计算与语言 · 计算机科学 2024-11-19 Joseph Liu , Mahesh Kumar Nandwana , Janne Pylkkönen , Hannes Heikinheimo , Morgan McGuire

With the increasing growth of social media, people have started relying heavily on the information shared therein to form opinions and make decisions. While such a reliance is motivation for a variety of parties to promote information, it…

计算与语言 · 计算机科学 2019-12-17 Rahul Radhakrishnan Iyer , Katia Sycara

Toxic language detection systems often falsely flag text that contains minority group mentions as toxic, as those groups are often the targets of online hate. Such over-reliance on spurious correlations also causes systems to struggle with…

计算与语言 · 计算机科学 2022-07-15 Thomas Hartvigsen , Saadia Gabriel , Hamid Palangi , Maarten Sap , Dipankar Ray , Ece Kamar

The use of abusive language online has become an increasingly pervasive problem that damages both individuals and society, with effects ranging from psychological harm right through to escalation to real-life violence and even death.…

计算与语言 · 计算机科学 2023-09-26 Mali Jin , Yida Mu , Diana Maynard , Kalina Bontcheva

Satire detection is essential for accurately extracting opinions from textual data and combating misinformation online. However, the lack of diverse corpora for satire leads to the problem of stylistic bias which impacts the models'…

计算与语言 · 计算机科学 2024-12-13 Asli Umay Ozturk , Recep Firat Cekinel , Pinar Karagoz

With surge in online platforms, there has been an upsurge in the user engagement on these platforms via comments and reactions. A large portion of such textual comments are abusive, rude and offensive to the audience. With machine learning…

计算与语言 · 计算机科学 2021-08-17 Ayush Kumar , Pratik Kumar

This work proposes a contextualised detection framework for implicitly hateful speech, implemented as a multi-agent system comprising a central Moderator Agent and dynamically constructed Community Agents representing specific demographic…

计算与语言 · 计算机科学 2026-01-28 Ewelina Gajewska , Katarzyna Budzynska , Jarosław A Chudziak