English
Related papers

Related papers: JANUS: A Lightweight Framework for Jailbreaking Te…

200 papers

Recently, the text-to-image diffusion model has gained considerable attention from the community due to its exceptional image generation capability. A representative model, Stable Diffusion, amassed more than 10 million users within just…

Cryptography and Security · Computer Science 2024-09-16 Chenyu Zhang , Mingwang Hu , Wenhui Li , Lanjun Wang

Despite rigorous safety alignment, Large Language Models (LLMs) remain vulnerable to jailbreak attacks. Existing black-box methods often rely on heuristic templates or exhaustive trials, lacking mechanistic interpretability and query…

Cryptography and Security · Computer Science 2026-05-19 Ziwei Wang , Jing Chen , Ruichao Liang , Zhi Wang , Yebo Feng , Ju Jia , Ruiying Du , Cong Wu , Yang Liu

While Large Language Models (LLMs) have achieved remarkable progress, they remain vulnerable to jailbreak attacks. Existing methods, primarily relying on discrete input optimization (e.g., GCG), often suffer from high computational costs…

Computation and Language · Computer Science 2026-01-09 Wenpeng Xing , Mohan Li , Chunqiang Hu , Haitao Xu , Ningyu Zhang , Bo Lin , Meng Han

Current text-to-image (T2I) synthesis diffusion models raise misuse concerns, particularly in creating prohibited or not-safe-for-work (NSFW) images. To address this, various safety mechanisms and red teaming attack methods are proposed to…

Cryptography and Security · Computer Science 2025-02-07 Pucheng Dang , Xing Hu , Dong Li , Rui Zhang , Qi Guo , Kaidi Xu

Text-to-image (T2I) models have gained significant popularity. Most of these are diffusion models with unique computational characteristics, distinct from both traditional small-scale ML models and large language models. They are highly…

Computer Vision and Pattern Recognition · Computer Science 2025-11-11 Shubham Agarwal , Subrata Mitra , Saud Iqbal

With the ability to generate high-quality images, text-to-image (T2I) models can be exploited for creating inappropriate content. To prevent misuse, existing safety measures are either based on text blacklists, which can be easily…

Computer Vision and Pattern Recognition · Computer Science 2024-08-20 Runtao Liu , Ashkan Khakzar , Jindong Gu , Qifeng Chen , Philip Torr , Fabio Pizzati

This paper explores a novel lightweight approach LightFair to achieve fair text-to-image diffusion models (T2I DMs) by addressing the adverse effects of the text encoder. Most existing methods either couple different parts of the diffusion…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Boyu Han , Qianqian Xu , Shilong Bao , Zhiyong Yang , Kangli Zi , Qingming Huang

Recent advances in text-to-image (T2I) diffusion models have enabled impressive generative capabilities, but they also raise significant safety concerns due to the potential to produce harmful or undesirable content. While concept erasure…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Nanxiang Jiang , Zhaoxin Fan , Enhan Kang , Daiheng Gao , Yun Zhou , Yanxia Chang , Zheng Zhu , Yeying Jin , Wenjun Wu

Large Language Models (LLMs) face threats from jailbreak prompts. Existing methods for defending against jailbreak attacks are primarily based on auxiliary models. These strategies, however, often require extensive data collection or…

Cryptography and Security · Computer Science 2025-11-21 Zhuoran Yang , Yanyong Zhang

Optimization-based jailbreaks typically adopt the Toxic-Continuation setting in large vision-language models (LVLMs), following the standard next-token prediction objective. In this setting, an adversarial image is optimized to make the…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Hee-Seon Kim , Minbeom Kim , Wonjun Lee , Kihyun Kim , Changick Kim

Text-to-image (T2I) diffusion models have become prominent tools for generating high-fidelity images from text prompts. However, when trained on unfiltered internet data, these models can produce unsafe, incorrect, or stylistically…

Computer Vision and Pattern Recognition · Computer Science 2024-09-11 Rohit Jena , Ali Taghibakhshi , Sahil Jain , Gerald Shen , Nima Tajbakhsh , Arash Vahdat

Text-to-Image (T2I) diffusion models have demonstrated strong generation ability, but their potential to generate unsafe content raises significant safety concerns. Existing inference-time defense methods typically perform category-agnostic…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Binhong Tan , Zhaoxin Wang , Handing Wang

Text-to-image (T2I) diffusion models have gained widespread application across various domains, demonstrating remarkable creative potential. However, the strong generalization capabilities of these models can inadvertently led they to…

Computer Vision and Pattern Recognition · Computer Science 2025-02-19 Die Chen , Zhiwen Li , Cen Chen , Xiaodan Li , Jinyan Ye

Recent developments in text-to-image models, particularly Stable Diffusion, have marked significant achievements in various applications. With these advancements, there are growing safety concerns about the vulnerability of the model that…

Computer Vision and Pattern Recognition · Computer Science 2024-01-18 Chenyu Zhang , Lanjun Wang , Anan Liu

In the past years, we have witnessed the remarkable success of Text-to-Image (T2I) models and their widespread use on the web. Extensive research in making T2I models produce hyper-realistic images has led to new concerns, such as…

Computer Vision and Pattern Recognition · Computer Science 2025-04-17 Muhammad Shahid Muneer , Simon S. Woo

By integrating language understanding with perceptual modalities such as images, multimodal large language models (MLLMs) constitute a critical substrate for modern AI systems, particularly intelligent agents operating in open and…

Cryptography and Security · Computer Science 2025-12-24 Songze Li , Jiameng Cheng , Yiming Li , Xiaojun Jia , Dacheng Tao

Text-to-image (T2I) models are widespread, but their limited safety guardrails expose end users to harmful content and potentially allow for model misuse. Current safety measures are typically limited to text-based filtering or concept…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Runtao Liu , I Chieh Chen , Jindong Gu , Jipeng Zhang , Renjie Pi , Qifeng Chen , Philip Torr , Ashkan Khakzar , Fabio Pizzati

Large Language Models (LLMs) are widely deployed in diverse real-world settings, yet remain vulnerable to jailbreaking, where prompt-based attacks bypass safety filters. We present THREAT (Targeted Harmful generation via Reframing and…

Cryptography and Security · Computer Science 2026-05-22 Shahnewaz Karim Sakib , Swati Kar , Anindya Bijoy Das

Recent advancements in text-to-image (T2I) models have unlocked a wide range of applications but also present significant risks, particularly in their potential to generate unsafe content. To mitigate this issue, researchers have developed…

Computer Vision and Pattern Recognition · Computer Science 2025-01-17 Yong-Hyun Park , Sangdoo Yun , Jin-Hwa Kim , Junho Kim , Geonhui Jang , Yonghyun Jeong , Junghyo Jo , Gayoung Lee

Vision-language models (VLMs) extend large language models (LLMs) with vision encoders, enabling text generation conditioned on both images and text. However, this multimodal integration expands the attack surface by exposing the model to…

Machine Learning · Computer Science 2026-02-03 Kaiyuan Cui , Yige Li , Yutao Wu , Xingjun Ma , Sarah Erfani , Christopher Leckie , Hanxun Huang