English
Related papers

Related papers: When Memory Becomes a Vulnerability: Towards Multi…

200 papers

Text-to-image generative models such as Stable Diffusion and DALL$\cdot$E raise many ethical concerns due to the generation of harmful images such as Not-Safe-for-Work (NSFW) ones. To address these ethical concerns, safety filters are often…

Machine Learning · Computer Science 2023-11-14 Yuchen Yang , Bo Hui , Haolin Yuan , Neil Gong , Yinzhi Cao

Large Language Models (LLMs) are widely deployed in diverse real-world settings, yet remain vulnerable to jailbreaking, where prompt-based attacks bypass safety filters. We present THREAT (Targeted Harmful generation via Reframing and…

Cryptography and Security · Computer Science 2026-05-22 Shahnewaz Karim Sakib , Swati Kar , Anindya Bijoy Das

Text-to-Image models may generate harmful content, such as pornographic images, particularly when unsafe prompts are submitted. To address this issue, safety filters are often added on top of text-to-image models, or the models themselves…

Cryptography and Security · Computer Science 2026-01-09 Zhengyuan Jiang , Yuepeng Hu , Yuchen Yang , Yinzhi Cao , Neil Zhenqiang Gong

Text-to-image (T2I) models can generate not-safe-for-work (NSFW) content, motivating multi-stage safety pipelines with both text and image filters. Newer LLM-based filters detect latent intent beyond keywords, making token-level…

Machine Learning · Computer Science 2026-05-26 Zixuan Chen , Hao Lin , Ke Xu , Xinghao Jiang , Tanfeng Sun

Text-based image generation models, such as Stable Diffusion and DALL-E 3, hold significant potential in content creation and publishing workflows, making them the focus in recent years. Despite their remarkable capability to generate…

Computation and Language · Computer Science 2025-06-04 Wenxuan Wang , Kuiyi Gao , Youliang Yuan , Jen-tse Huang , Qiuzhi Liu , Shuai Wang , Wenxiang Jiao , Zhaopeng Tu

The rapid development of generative artificial intelligence has made text to video models essential for building future multimodal world simulators. However, these models remain vulnerable to jailbreak attacks, where specially crafted…

Cryptography and Security · Computer Science 2025-04-29 Siyuan Liang , Jiayang Liu , Jiecheng Zhai , Tianmeng Fang , Rongcheng Tu , Aishan Liu , Xiaochun Cao , Dacheng Tao

Text-to-image (T2I) models have significantly advanced in producing high-quality images. However, such models have the ability to generate images containing not-safe-for-work (NSFW) content, such as pornography, violence, political content,…

Cryptography and Security · Computer Science 2025-05-15 Longtian Wang , Xiaofei Xie , Tianlin Li , Yuhan Zhi , Chao Shen

Recent advances in image generation models (IGMs), particularly diffusion-based architectures such as Stable Diffusion (SD), have markedly enhanced the quality and diversity of AI-generated visual content. However, their generative…

Computer Vision and Pattern Recognition · Computer Science 2026-02-17 Renyang Liu , Guanlin Li , Tianwei Zhang , See-Kiong Ng

The widespread adoption of thinking mode in large language models (LLMs) has significantly enhanced complex task processing capabilities while introducing new security risks. When subjected to jailbreak attacks, the step-by-step reasoning…

Cryptography and Security · Computer Science 2026-03-12 Fan Yang

Text-to-Image (T2I) diffusion models have demonstrated strong generation ability, but their potential to generate unsafe content raises significant safety concerns. Existing inference-time defense methods typically perform category-agnostic…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Binhong Tan , Zhaoxin Wang , Handing Wang

Unified Multimodal understanding and generation Models (UMMs) have demonstrated remarkable capabilities in both understanding and generation tasks. However, we identify a vulnerability arising from the generation-understanding coupling in…

Artificial Intelligence · Computer Science 2025-10-01 Shaoxiong Guo , Tianyi Du , Lijun Li , Yuyao Wu , Jie Li , Jing Shao

While text-to-image synthesis currently enjoys great popularity among researchers and the general public, the security of these models has been neglected so far. Many text-guided image generation models rely on pre-trained text encoders…

Machine Learning · Computer Science 2023-08-10 Lukas Struppek , Dominik Hintersdorf , Kristian Kersting

Text-to-Image generation models have revolutionized the artwork design process and enabled anyone to create high-quality images by entering text descriptions called prompts. Creating a high-quality prompt that consists of a subject and…

Cryptography and Security · Computer Science 2024-04-16 Xinyue Shen , Yiting Qu , Michael Backes , Yang Zhang

Text-to-image (T2I) models have demonstrated remarkable generative capabilities but remain vulnerable to producing not-safe-for-work (NSFW) content, such as violent or explicit imagery. While recent moderation efforts have introduced soft…

Computer Vision and Pattern Recognition · Computer Science 2025-08-15 Zonglei Jing , Xiao Yang , Xiaoqian Li , Siyuan Liang , Aishan Liu , Mingchuan Zhang , Xianglong Liu

In the past few years, Language Models (LMs) have shown par-human capabilities in several domains. Despite their practical applications and exceeding user consumption, they are susceptible to jailbreaks when malicious input exploits the…

Computation and Language · Computer Science 2025-04-18 Charlotte Siska , Anush Sankaran

The generative AI revolution in recent years has been spurred by an expansion in compute power and data quantity, which together enable extensive pre-training of powerful text-to-image (T2I) models. With their greater capabilities to…

Large (vision-)language models exhibit remarkable capability but remain highly susceptible to jailbreaking. Existing safety training approaches aim to have the model learn a refusal boundary between safe and unsafe, based on the user's…

Cryptography and Security · Computer Science 2026-04-28 Xinhe Wang , Katia Sycara , Yaqi Xie

In recent years, text-to-image (T2I) generation models have made significant progress in generating high-quality images that align with text descriptions. However, these models also face the risk of unsafe generation, potentially producing…

Cryptography and Security · Computer Science 2025-04-16 Huming Qiu , Guanxu Chen , Mi Zhang , Xiaohan Zhang , Xiaoyu You , Min Yang

The increasing sophistication of large vision-language models (LVLMs) has been accompanied by advances in safety alignment mechanisms designed to prevent harmful content generation. However, these defenses remain vulnerable to sophisticated…

Cryptography and Security · Computer Science 2026-04-09 Quanchen Zou , Zonghao Ying , Moyang Chen , Wenzhuo Xu , Yisong Xiao , Yakai Li , Deyue Zhang , Dongdong Yang , Zhao Liu , Xiangzheng Zhang

Text-to-Image (T2I) models, represented by DALL$\cdot$E and Midjourney, have gained huge popularity for creating realistic images. The quality of these images relies on the carefully engineered prompts, which have become valuable…

Cryptography and Security · Computer Science 2026-01-22 Shiqian Zhao , Chong Wang , Yiming Li , Yihao Huang , Wenjie Qu , Siew-Kei Lam , Yi Xie , Kangjie Chen , Jie Zhang , Tianwei Zhang