中文
相关论文

相关论文: Exploiting Explainability to Design Adversarial At…

200 篇论文

Large Language Models (LLMs) are increasingly vulnerable to adversarial attacks that can subtly manipulate their outputs. While various defense mechanisms have been proposed, many operate as black boxes, lacking transparency in their…

密码学与安全 · 计算机科学 2025-11-19 Shaowei Guan , Yu Zhai , Zhengyu Zhang , Yanze Wang , Hin Chi Kwok

Hate speech is one type of harmful online content which directly attacks or promotes hate towards a group or an individual member based on their actual or perceived aspects of identity, such as ethnicity, religion, and sexual orientation.…

计算与语言 · 计算机科学 2021-02-18 Wenjie Yin , Arkaitz Zubiaga

With the spread of social networks and their unfortunate use for hate speech, automatic detection of the latter has become a pressing problem. In this paper, we reproduce seven state-of-the-art hate speech detection models from prior work,…

计算与语言 · 计算机科学 2018-11-06 Tommi Gröndahl , Luca Pajola , Mika Juuti , Mauro Conti , N. Asokan

The proliferation of multimodal content on social media presents significant challenges in understanding and moderating complex, context-dependent issues such as misinformation, hate speech, and propaganda. While efforts have been made to…

计算与语言 · 计算机科学 2026-03-03 Mohamed Bayan Kmainasi , Abul Hasnat , Md Arid Hasan , Ali Ezzat Shahroor , Firoj Alam

With the increasing diversity of use cases of large language models, a more informative treatment of texts seems necessary. An argumentative analysis could foster a more reasoned usage of chatbots, text completion mechanisms or other…

计算与语言 · 计算机科学 2023-06-06 Damián Furman , Pablo Torres , José A. Rodríguez , Diego Letzen , Vanina Martínez , Laura Alonso Alemany

Social media platforms are deploying machine learning based offensive language classification systems to combat hateful, racist, and other forms of offensive speech at scale. However, despite their real-world deployment, we do not yet…

计算与语言 · 计算机科学 2022-03-23 Jonathan Rusert , Zubair Shafiq , Padmini Srinivasan

In the evolving landscape of online communication, hate speech detection remains a formidable challenge, further compounded by the diversity of digital platforms. This study investigates the effectiveness and adaptability of pre-trained and…

计算与语言 · 计算机科学 2025-05-01 Ahmad Nasir , Aadish Sharma , Kokil Jaidka , Saifuddin Ahmed

Hate speech detection on online social networks has become one of the emerging hot topics in recent years. With the broad spread and fast propagation speed across online social networks, hate speech makes significant impacts on society by…

计算与语言 · 计算机科学 2024-09-26 Guanyi Mou , Pengyi Ye , Kyumin Lee

Over the past decade, there has been extensive research aimed at enhancing the robustness of neural networks, yet this problem remains vastly unsolved. Here, one major impediment has been the overestimation of the robustness of new defense…

人工智能 · 计算机科学 2023-10-31 Leo Schwinn , David Dobre , Stephan Günnemann , Gauthier Gidel

Hate speech is a harmful form of online expression, often manifesting as derogatory posts. It is a significant risk in digital environments. With the rise of Large Language Models (LLMs), there is concern about their potential to replicate…

计算与语言 · 计算机科学 2025-06-10 Paloma Piot , Javier Parapar

Automated hate speech detection is an important tool in combating the spread of hate speech, particularly in social media. Numerous methods have been developed for the task, including a recent proliferation of deep-learning based…

计算与语言 · 计算机科学 2023-12-08 Jitendra Singh Malik , Hezhe Qiao , Guansong Pang , Anton van den Hengel

Social media platforms enable instant and ubiquitous connectivity and are essential to social interaction and communication in our technological society. Apart from its advantages, these platforms have given rise to negative behaviors in…

社会与信息网络 · 计算机科学 2025-05-08 Silvia García-Méndez , Francisco De Arriba-Pérez

The prevalence of offensive content on the internet, encompassing hate speech and cyberbullying, is a pervasive issue worldwide. Consequently, it has garnered significant attention from the machine learning (ML) and natural language…

计算与语言 · 计算机科学 2024-07-29 Alphaeus Dmonte , Tejas Arya , Tharindu Ranasinghe , Marcos Zampieri

Hate speech is increasingly prevalent online, and its negative outcomes include increased prejudice, extremism, and even offline hate crime. Automatic detection of online hate speech can help us to better understand these impacts. However,…

计算与语言 · 计算机科学 2021-02-10 John D Gallacher

The growing use of social media has led to the development of several Machine Learning (ML) and Natural Language Processing(NLP) tools to process the unprecedented amount of social media content to make actionable decisions. However, these…

计算与语言 · 计算机科学 2021-10-28 Izzat Alsmadi , Kashif Ahmad , Mahmoud Nazzal , Firoj Alam , Ala Al-Fuqaha , Abdallah Khreishah , Abdulelah Algosaibi

The age of social media is flooded with Internet memes, necessitating a clear grasp and effective identification of harmful ones. This task presents a significant challenge due to the implicit meaning embedded in memes, which is not…

计算与语言 · 计算机科学 2024-01-25 Hongzhan Lin , Ziyang Luo , Wei Gao , Jing Ma , Bo Wang , Ruichao Yang

The spread of hate speech on social media space is currently a serious issue. The undemanding access to the enormous amount of information being generated on these platforms has led people to post and react with toxic content that…

计算与语言 · 计算机科学 2022-09-13 Abhishek Velankar , Hrushikesh Patil , Raviraj Joshi

The assessment of legal problems requires the consideration of a specific legal system and its levels of abstraction, from constitutional law to statutory law to case law. The extent to which Large Language Models (LLMs) internalize such…

计算与语言 · 计算机科学 2025-06-04 Florian Ludwig , Torsten Zesch , Frederike Zufall

Detecting hateful content is a challenging and important problem. Automated tools, like machine-learning models, can help, but they require continuous training to adapt to the ever-changing landscape of social media. In this work, we…

计算与语言 · 计算机科学 2025-11-06 Jay Patel , Hrudayangam Mehta , Jeremy Blackburn

User generated text on social media often suffers from a lot of undesired characteristics including hatespeech, abusive language, insults etc. that are targeted to attack or abuse a specific group of people. Often such text is written…

计算与语言 · 计算机科学 2019-10-03 Sravan Babu Bodapati , Spandana Gella , Kasturi Bhattacharjee , Yaser Al-Onaizan