中文
相关论文

相关论文: Cyberbullying Classifiers are Sensitive to Model-A…

200 篇论文

Fake news detection models are critical to countering disinformation but can be manipulated through adversarial attacks. In this position paper, we analyze how an attacker can compromise the performance of an online learning detector on…

机器学习 · 计算机科学 2024-01-05 Federico Siciliano , Luca Maiano , Lorenzo Papa , Federica Baccini , Irene Amerini , Fabrizio Silvestri

Augmenting toxic language data in a controllable and class-specific manner is crucial for improving robustness in toxicity classification, yet remains challenging due to limited supervision and distributional skew. We propose ToxiGAN, a…

计算与语言 · 计算机科学 2026-01-07 Peiran Li , Jan Fillies , Adrian Paschke

Uses of pejorative expressions can be benign or actively empowering. When models for abuse detection misclassify these expressions as derogatory, they inadvertently censor productive conversations held by marginalized groups. One way to…

计算与语言 · 计算机科学 2022-06-20 Jana Kurrek , Haji Mohammad Saleem , Derek Ruths

Social media continues to have an impact on the trajectory of humanity. However, its introduction has also weaponized keyboards, allowing the abusive language normally reserved for in-person bullying to jump onto the screen, i.e.,…

机器学习 · 计算机科学 2024-09-20 Muhammad Arslan , Manuel Sandoval Madrigal , Mohammed Abuhamad , Deborah L. Hall , Yasin N. Silva

Detecting cyberbullying on social media remains a critical challenge due to its subtle and varied expressions. This study investigates whether integrating aggression detection as an auxiliary task within a unified training framework can…

计算与语言 · 计算机科学 2025-08-25 Aisha Saeid , Anu Sabu , Girish A. Koushik , Ferrante Neri , Diptesh Kanojia

Technological advancements have resulted in an exponential increase in the use of online social networks (OSNs) worldwide. While online social networks provide a great communication medium, they also increase the user's exposure to…

计算机与社会 · 计算机科学 2024-01-09 Sylvia W Azumah , Nelly Elsayed , Zag ElSayed , Murat Ozer

Large language models produce human-like text that drive a growing number of applications. However, recent literature and, increasingly, real world observations, have demonstrated that these models can generate language that is toxic,…

Although automated harmful content detection systems are frequently used to monitor online platforms, moderators and end users frequently cannot understand the logic underlying their predictions. While recent studies have focused on…

计算与语言 · 计算机科学 2026-03-20 Trishita Dhara , Siddhesh Sheth

Cyberbullying is a prevalent social problem that inflicts detrimental consequences to the health and safety of victims such as psychological distress, anti-social behaviour, and suicide. The automation of cyberbullying detection is a recent…

计算与语言 · 计算机科学 2020-10-26 Gathika Ratnayaka , Thushari Atapattu , Mahen Herath , Georgia Zhang , Katrina Falkner

The detection of online cyberbullying has seen an increase in societal importance, popularity in research, and available open data. Nevertheless, while computational power and affordability of resources continue to increase, the access…

Social media platforms are plagued by harmful content such as hate speech, misinformation, and extremist rhetoric. Machine learning (ML) models are widely adopted to detect such content; however, they remain highly vulnerable to adversarial…

机器学习 · 计算机科学 2025-12-30 Yidong Chai , Yi Liu , Mohammadreza Ebrahimi , Weifeng Li , Balaji Padmanabhan

Classifiers tend to propagate biases present in the data on which they are trained. Hence, it is important to understand how the demographic identities of the annotators of comments affect the fairness of the resulting model. In this paper,…

计算与语言 · 计算机科学 2021-06-07 Elizabeth Excell , Noura Al Moubayed

We present two categories of model-agnostic adversarial strategies that reveal the weaknesses of several generative, task-oriented dialogue models: Should-Not-Change strategies that evaluate over-sensitivity to small and…

计算与语言 · 计算机科学 2018-09-07 Tong Niu , Mohit Bansal

Machine learning classifiers are known to be vulnerable to inputs maliciously constructed by adversaries to force misclassification. Such adversarial examples have been extensively studied in the context of computer vision applications. In…

机器学习 · 计算机科学 2017-02-09 Sandy Huang , Nicolas Papernot , Ian Goodfellow , Yan Duan , Pieter Abbeel

When trained on large, unfiltered crawls from the internet, language models pick up and reproduce all kinds of undesirable biases that can be found in the data: they often generate racist, sexist, violent or otherwise toxic language. As…

计算与语言 · 计算机科学 2021-09-10 Timo Schick , Sahana Udupa , Hinrich Schütze

As machine learning systems become more widely used, especially for safety critical applications, there is a growing need to ensure that these systems behave as intended, even in the face of adversarial examples. Adversarial examples are…

计算与语言 · 计算机科学 2024-08-19 Anahita Samadi , Allison Sullivan

The proliferation of social media platforms has led to an increase in the spread of hate speech, particularly targeting vulnerable communities. Unfortunately, existing methods for automatically identifying and blocking toxic language rely…

计算与语言 · 计算机科学 2025-02-24 Shiza Ali , Jeremy Blackburn , Gianluca Stringhini

Cyberbullying is a significant concern intricately linked to technology that can find resolution through technological means. Despite its prevalence, technology also provides solutions to mitigate cyberbullying. To address growing concerns…

机器学习 · 计算机科学 2024-06-27 Sylvia Worlali Azumah , Nelly Elsayed , Zag ElSayed , Murat Ozer , Amanda La Guardia

Model distillation has become essential for creating smaller, deployable language models that retain larger system capabilities. However, widespread deployment raises concerns about resilience to adversarial manipulation. This paper…

机器学习 · 计算机科学 2025-10-17 Harsh Chaudhari , Jamie Hayes , Matthew Jagielski , Ilia Shumailov , Milad Nasr , Alina Oprea

With the recent rise of toxicity in online conversations on social media platforms, using modern machine learning algorithms for toxic comment detection has become a central focus of many online applications. Researchers and companies have…

人工智能 · 计算机科学 2020-03-30 Ameya Vaidya , Feng Mai , Yue Ning