English
Related papers

Related papers: Exploiting Explainability to Design Adversarial At…

200 papers

The proliferation of social media platforms has led to an increase in the spread of hate speech, particularly targeting vulnerable communities. Unfortunately, existing methods for automatically identifying and blocking toxic language rely…

Computation and Language · Computer Science 2025-02-24 Shiza Ali , Jeremy Blackburn , Gianluca Stringhini

Transformer-based text classifiers such as BERT, RoBERTa, T5, and GPT have shown strong performance in natural language processing tasks but remain vulnerable to adversarial examples. These vulnerabilities raise significant security…

Computation and Language · Computer Science 2025-10-27 Bushra Sabir , Yansong Gao , Alsharif Abuadbba , M. Ali Babar

Optimization of offensive content moderation models for different types of hateful messages is typically achieved through continued pre-training or fine-tuning on new hate speech benchmarks. However, existing benchmarks mainly address…

Computation and Language · Computer Science 2026-04-07 Irina Proskurina , Marc-Antoine Carpentier , Julien Velcin

Training robust deep learning models for down-stream tasks is a critical challenge. Research has shown that down-stream models can be easily fooled with adversarial inputs that look like the training data, but slightly perturbed, in a way…

Machine Learning · Computer Science 2021-01-19 Mahmoud Hossam , Trung Le , He Zhao , Dinh Phung

Detecting and classifying instances of hate in social media text has been a problem of interest in Natural Language Processing in the recent years. Our work leverages state of the art Transformer language models to identify hate speech in a…

Computation and Language · Computer Science 2021-01-12 Sayar Ghosh Roy , Ujwal Narayan , Tathagata Raha , Zubair Abid , Vasudeva Varma

Hate speech in social media is a growing phenomenon, and detecting such toxic content has recently gained significant traction in the research community. Existing studies have explored fine-tuning language models (LMs) to perform hate…

Computation and Language · Computer Science 2023-03-07 Md Rabiul Awal , Roy Ka-Wei Lee , Eshaan Tanwar , Tanmay Garg , Tanmoy Chakraborty

Offensive or antagonistic language targeted at individuals and social groups based on their personal characteristics (also known as cyber hate speech or cyberhate) has been frequently posted and widely circulated viathe World Wide Web. This…

Computation and Language · Computer Science 2018-03-09 Wafa Alorainy , Pete Burnap , Han Liu , Matthew Williams

Automatic hate speech detection in online social networks is an important open problem in Natural Language Processing (NLP). Hate speech is a multidimensional issue, strongly dependant on language and cultural factors. Despite its…

Computation and Language · Computer Science 2021-05-03 Aymé Arango , Jorge Pérez , Barbara Poblete

Recently efforts have been made by social media platforms as well as researchers to detect hateful or toxic language using large language models. However, none of these works aim to use explanation, additional context and victim community…

Computation and Language · Computer Science 2023-10-31 Sarthak Roy , Ashish Harshavardhan , Animesh Mukherjee , Punyajoy Saha

Hateful memes are an emerging method of spreading hate on the internet, relying on both images and text to convey a hateful message. We take an interpretable approach to hateful meme detection, using machine learning and simple heuristics…

Machine Learning · Computer Science 2021-08-24 Tanvi Deshpande , Nitya Mani

The detection of offensive, hateful and profane language has become a critical challenge since many users in social networks are exposed to cyberbullying activities on a daily basis. In this paper, we present an analysis of combining…

Computation and Language · Computer Science 2021-12-10 Sherzod Hakimov , Ralph Ewerth

This work investigates the potential of undermining both fairness and detection performance in abusive language detection. In a dynamic and complex digital world, it is crucial to investigate the vulnerabilities of these detection models to…

Computation and Language · Computer Science 2023-12-07 Yueqing Liang , Lu Cheng , Ali Payani , Kai Shu

With the recent advancements in machine learning (ML), numerous ML-based approaches have been extensively applied in software analytics tasks to streamline software development and maintenance processes. Nevertheless, studies indicate that…

Software Engineering · Computer Science 2025-07-15 MD Abdul Awal , Mrigank Rochan , Chanchal K. Roy

Online hate speech is an important issue that breaks the cohesiveness of online social communities and even raises public safety concerns in our societies. Motivated by this rising issue, researchers have developed many traditional machine…

Computation and Language · Computer Science 2021-03-23 Rui Cao , Roy Ka-Wei Lee , Tuan-Anh Hoang

In the current context where online platforms have been effectively weaponized in a variety of geo-political events and social issues, Internet memes make fair content moderation at scale even more difficult. Existing work on meme…

Artificial Intelligence · Computer Science 2023-04-10 Abhinav Kumar Thakur , Filip Ilievski , Hông-Ân Sandlin , Zhivar Sourati , Luca Luceri , Riccardo Tommasini , Alain Mermoud

The rapid evolution of social media has provided enhanced communication channels for individuals to create online content, enabling them to express their thoughts and opinions. Multimodal memes, often utilized for playful or humorous…

Computer Vision and Pattern Recognition · Computer Science 2025-05-02 Minh-Hao Van , Xintao Wu

Hate speech, offensive language, aggression, racism, sexism, and other abusive language are common phenomena in social media. There is a need for Artificial Intelligence(AI)based intervention which can filter hate content at scale. Most…

Computation and Language · Computer Science 2024-11-13 Prashant Kapil , Asif Ekbal

Large Language Models (LLMs) are swiftly advancing in architecture and capability, and as they integrate more deeply into complex systems, the urgency to scrutinize their security properties grows. This paper surveys research in the…

Computation and Language · Computer Science 2023-10-18 Erfan Shayegani , Md Abdullah Al Mamun , Yu Fu , Pedram Zaree , Yue Dong , Nael Abu-Ghazaleh

Digital platforms have an ever-expanding user base, and act as a hub for communication, business, and connectivity. However, this has also allowed for the spread of hate speech and misogyny. Artificial intelligence models have emerged as an…

Artificial Intelligence · Computer Science 2026-01-14 Sargam Yadav , Abhishek Kaushik , Kevin Mc Daid

Hate speech and misinformation, spread over social networking services (SNS) such as Facebook and Twitter, have inflamed ethnic and political violence in countries across the globe. We argue that there is limited research on this problem…

Social and Information Networks · Computer Science 2022-08-23 Cuong Nguyen , Daniel Nkemelu , Ankit Mehta , Michael Best