中文
相关论文

相关论文: Hate in Plain Sight: On the Risks of Moderating AI…

200 篇论文

Hate speech online targets individuals or groups based on identity attributes and spreads rapidly, posing serious social risks. Memes, which combine images and text, have emerged as a nuanced vehicle for disseminating hate speech, often…

多智能体系统 · 计算机科学 2026-03-26 Rui Xing , Qi Chai , Jie Ma , Jing Tao , Pinghui Wang , Shuming Zhang , Xinping Wang , Hao Wang

Text-to-image diffusion models have achieved widespread popularity due to their unprecedented image generation capability. In particular, their ability to synthesize and modify human faces has spurred research into using generated face…

计算机视觉与模式识别 · 计算机科学 2023-12-22 Harrison Rosenberg , Shimaa Ahmed , Guruprasad V Ramesh , Ramya Korlakai Vinayak , Kassem Fawaz

This research introduces a novel approach to textual and multimodal Hate Speech Detection (HSD), using Large Language Models (LLMs) as dynamic knowledge bases to generate background context and incorporate it into the input of HSD…

计算与语言 · 计算机科学 2025-10-20 Joshua Wolfe Brook , Ilia Markov

From uncertainty quantification to real-world object detection, we recognize the importance of machine learning algorithms, particularly in safety-critical domains such as autonomous driving or medical diagnostics. In machine learning,…

计算机视觉与模式识别 · 计算机科学 2025-05-29 Carina Newen , Luca Hinkamp , Maria Ntonti , Emmanuel Müller

Recently, large language models (LLMs) have taken the spotlight in natural language processing. Further, integrating LLMs with vision enables the users to explore more emergent abilities in multimodality. Visual language models (VLMs), such…

计算与语言 · 计算机科学 2023-11-14 Minh-Hao Van , Xintao Wu

Hate speech is increasingly prevalent online, and its negative outcomes include increased prejudice, extremism, and even offline hate crime. Automatic detection of online hate speech can help us to better understand these impacts. However,…

计算与语言 · 计算机科学 2021-02-10 John D Gallacher

Hate speech is a harmful form of online expression, often manifesting as derogatory posts. It is a significant risk in digital environments. With the rise of Large Language Models (LLMs), there is concern about their potential to replicate…

计算与语言 · 计算机科学 2025-06-10 Paloma Piot , Javier Parapar

Hateful memes often require compositional multimodal reasoning: the image and text may appear benign in isolation, yet their interaction conveys harmful intent. Although thinking-based multimodal large language models (MLLMs) have recently…

计算与语言 · 计算机科学 2026-03-03 Mohamed Bayan Kmainasi , Mucahid Kutlu , Ali Ezzat Shahroor , Abul Hasnat , Firoj Alam

Deepfake or synthetic images produced using deep generative models pose serious risks to online platforms. This has triggered several research efforts to accurately detect deepfake images, achieving excellent performance on publicly…

Memes on the Internet are often harmless and sometimes amusing. However, by using certain types of images, text, or combinations of both, the seemingly harmless meme becomes a multimodal type of hate speech -- a hateful meme. The Hateful…

人工智能 · 计算机科学 2020-12-25 Riza Velioglu , Jewgeni Rose

Hate speech is one type of harmful online content which directly attacks or promotes hate towards a group or an individual member based on their actual or perceived aspects of identity, such as ethnicity, religion, and sexual orientation.…

计算与语言 · 计算机科学 2021-02-18 Wenjie Yin , Arkaitz Zubiaga

Diffusion-based models have gained significant popularity for text-to-image generation due to their exceptional image-generation capabilities. A risk with these models is the potential generation of inappropriate content, such as biased or…

计算机视觉与模式识别 · 计算机科学 2024-03-29 Hang Li , Chengzhi Shen , Philip Torr , Volker Tresp , Jindong Gu

Current image generation models can effortlessly produce high-quality, highly realistic images, but this also increases the risk of misuse. In various Text-to-Image or Image-to-Image tasks, attackers can generate a series of images…

计算机视觉与模式识别 · 计算机科学 2025-05-01 Hao Cheng , Erjia Xiao , Jiayan Yang , Jiahang Cao , Qiang Zhang , Jize Zhang , Kaidi Xu , Jindong Gu , Renjing Xu

Warning: this paper contains content that may be offensive or upsetting Hate speech moderation on global platforms poses unique challenges due to the multimodal and multilingual nature of content, along with the varying cultural…

计算与语言 · 计算机科学 2025-02-18 Minh Duc Bui , Katharina von der Wense , Anne Lauscher

The evolution of Artificial Intelligence Generated Contents (AIGCs) is advancing towards higher quality. The growing interactions with AIGCs present a new challenge to the data-driven AI community: While AI-generated contents have played a…

计算机视觉与模式识别 · 计算机科学 2024-09-04 Yifei Gao , Jiaqi Wang , Zhiyu Lin , Jitao Sang

Although pretrained large language models (PLMs) have achieved state-of-the-art on many natural language processing (NLP) tasks, they lack an understanding of subtle expressions of implicit hate speech. Various attempts have been made to…

计算与语言 · 计算机科学 2026-03-04 Sarah Masud , Ashutosh Bajpai , Tanmoy Chakraborty

Hate speech detection is a critical problem in social media platforms, being often accused for enabling the spread of hatred and igniting physical violence. Hate speech detection requires overwhelming resources including high-performance…

计算与语言 · 计算机科学 2020-05-14 Tomer Wullach , Amir Adler , Einat Minkov

Vision-language models (VLMs) exhibit a systematic bias when confronted with classic optical illusions: they overwhelmingly predict the illusion as "real" regardless of whether the image has been counterfactually modified. We present a…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Xuesong Wang , Harry Wang

State-of-the-art Diffusion Models (DMs) produce highly realistic images. While prior work has successfully mitigated Not Safe For Work (NSFW) content in the visual domain, we identify a novel threat: the generation of NSFW text embedded…

计算机视觉与模式识别 · 计算机科学 2026-01-16 Aditya Kumar , Tom Blanchard , Adam Dziedzic , Franziska Boenisch

Automatic hate speech detection using deep neural models is hampered by the scarcity of labeled datasets, leading to poor generalization. To mitigate this problem, generative AI has been utilized to generate large amounts of synthetic hate…

计算与语言 · 计算机科学 2023-11-17 Sagi Pendzel , Tomer Wullach , Amir Adler , Einat Minkov