English
Related papers

Related papers: Not-in-Perspective: Towards Shielding Google's Per…

200 papers

Detecting "toxic" language in internet content is a pressing social and technical challenge. In this work, we focus on PERSPECTIVE from Jigsaw, a state-of-the-art tool that promises to score the "toxicity" of text, with a recent model…

Computation and Language · Computer Science 2023-01-06 Lorena Piedras , Lucas Rosenblatt , Julia Wilkins

The censorship of toxic comments is often left to the judgment of imperfect models. Perspective API, a creation of Google technology incubator Jigsaw, is perhaps the most widely used toxicity classifier in industry; the model is employed by…

Computers and Society · Computer Science 2020-07-01 Elizabeth Reichert , Helen Qiu , Jasmine Bayrooti

With machine learning models being used for more sensitive applications, we rely on interpretability methods to prove that no discriminating attributes were used for classification. A potential concern is the so-called "fair-washing" -…

Machine Learning · Computer Science 2020-07-14 Laura Rieger , Lars Kai Hansen

Recent studies have shown that deep neural networks are vulnerable to intentionally crafted adversarial examples, and various methods have been proposed to defend against adversarial word-substitution attacks for neural NLP models. However,…

Computation and Language · Computer Science 2021-10-07 Zongyi Li , Jianhan Xu , Jiehang Zeng , Linyang Li , Xiaoqing Zheng , Qi Zhang , Kai-Wei Chang , Cho-Jui Hsieh

Safety classifiers are critical in mitigating toxicity on online forums such as social media and in chatbots. Still, they continue to be vulnerable to emergent, and often innumerable, adversarial attacks. Traditional automated adversarial…

Computation and Language · Computer Science 2024-06-26 Yash Kumar Lal , Preethi Lahoti , Aradhana Sinha , Yao Qin , Ananth Balashankar

The spectacular expansion of the Internet has led to the development of a new research problem in the field of natural language processing: automatic toxic comment detection, since many countries prohibit hate speech in public media. There…

Machine Learning · Computer Science 2020-09-18 Ashwin Geet D'Sa , Irina Illina , Dominique Fohr

The rise of foundation models fine-tuned on human feedback from potentially untrusted users has increased the risk of adversarial data poisoning, necessitating the study of robustness of learning algorithms against such attacks. Existing…

Machine Learning · Computer Science 2025-02-25 Avinandan Bose , Laurent Lessard , Maryam Fazel , Krishnamurthy Dj Dvijotham

Cyberbullying is a significant concern intricately linked to technology that can find resolution through technological means. Despite its prevalence, technology also provides solutions to mitigate cyberbullying. To address growing concerns…

Machine Learning · Computer Science 2024-06-27 Sylvia Worlali Azumah , Nelly Elsayed , Zag ElSayed , Murat Ozer , Amanda La Guardia

With the development of large language models (LLMs), detecting whether text is generated by a machine becomes increasingly challenging in the face of malicious use cases like the spread of false information, protection of intellectual…

Computation and Language · Computer Science 2024-04-03 Ying Zhou , Ben He , Le Sun

This paper investigates the use of machine learning models for the classification of unhealthy online conversations containing one or more forms of subtler abuse, such as hostility, sarcasm, and generalization. We leveraged a public dataset…

Computation and Language · Computer Science 2022-01-28 Shlok Gilda , Mirela Silva , Luiz Giovanini , Daniela Oliveira

Social media platforms are deploying machine learning based offensive language classification systems to combat hateful, racist, and other forms of offensive speech at scale. However, despite their real-world deployment, we do not yet…

Computation and Language · Computer Science 2022-03-23 Jonathan Rusert , Zubair Shafiq , Padmini Srinivasan

Recently, advanced NLP models have seen a surge in the usage of various applications. This raises the security threats of the released models. In addition to the clean models' unintentional weaknesses, {\em i.e.,} adversarial attacks, the…

Computation and Language · Computer Science 2021-01-18 Lichao Sun

Social media platforms enable instant and ubiquitous connectivity and are essential to social interaction and communication in our technological society. Apart from its advantages, these platforms have given rise to negative behaviors in…

Social and Information Networks · Computer Science 2025-05-08 Silvia García-Méndez , Francisco De Arriba-Pérez

Several years of research have shown that machine-learning systems are vulnerable to adversarial examples, both in theory and in practice. Until now, such attacks have primarily targeted visual models, exploiting the gap between human and…

Computation and Language · Computer Science 2021-12-14 Nicholas Boucher , Ilia Shumailov , Ross Anderson , Nicolas Papernot

The evolution of digital communication systems and the designs of online platforms have inadvertently facilitated the subconscious propagation of toxic behavior. Giving rise to reactive responses to toxic behavior. Toxicity in online…

Computers and Society · Computer Science 2025-10-01 Smita Khapre , Melkamu Abay Mersha , Hassan Shakil , Jonali Baruah , Jugal Kalita

The landscape of available textual adversarial attacks keeps growing, posing severe threats and raising concerns regarding the deep NLP system's integrity. However, the crucial problem of defending against malicious attacks has only drawn…

Computation and Language · Computer Science 2023-10-24 Pierre Colombo , Marine Picot , Nathan Noiry , Guillaume Staerman , Pablo Piantanida

Automatic detection of toxic language plays an essential role in protecting social media users, especially minority groups, from verbal abuse. However, biases toward some attributes, including gender, race, and dialect, exist in most…

Computation and Language · Computer Science 2021-06-15 Yung-Sung Chuang , Mingye Gao , Hongyin Luo , James Glass , Hung-yi Lee , Yun-Nung Chen , Shang-Wen Li

Social media platforms are plagued by harmful content such as hate speech, misinformation, and extremist rhetoric. Machine learning (ML) models are widely adopted to detect such content; however, they remain highly vulnerable to adversarial…

Machine Learning · Computer Science 2025-12-30 Yidong Chai , Yi Liu , Mohammadreza Ebrahimi , Weifeng Li , Balaji Padmanabhan

Flood of information is produced in a daily basis through the global Internet usage arising from the on-line interactive communications among users. While this situation contributes significantly to the quality of human life, unfortunately…

Computation and Language · Computer Science 2024-06-04 Spiros V. Georgakopoulos , Sotiris K. Tasoulis , Aristidis G. Vrahatis , Vassilis P. Plagianakos

A limited amount of studies investigates the role of model-agnostic adversarial behavior in toxic content classification. As toxicity classifiers predominantly rely on lexical cues, (deliberately) creative and evolving language-use can be…

Computation and Language · Computer Science 2022-01-19 Chris Emmery , Ákos Kádár , Grzegorz Chrupała , Walter Daelemans