English
Related papers

Related papers: Automatic Jailbreaking of the Text-to-Image Genera…

200 papers

Despite significant advancements in alignment and content moderation, large language models (LLMs) and text-to-image (T2I) systems remain vulnerable to prompt-based attacks known as jailbreaks. Unlike traditional adversarial examples…

Computer Vision and Pattern Recognition · Computer Science 2025-07-30 Ahmed B Mustafa , Zihan Ye , Yang Lu , Michael P Pound , Shreyank N Gowda

As large language models (LLMs) advance, ensuring AI safety and alignment is paramount. One popular approach is prompt guards, lightweight mechanisms designed to filter malicious queries while being easy to implement and update. In this…

Machine Learning · Computer Science 2025-10-08 Jaiden Fairoze , Sanjam Garg , Keewoo Lee , Mingyuan Wang

Recent advancements in generative AI have enabled ubiquitous access to large language models (LLMs). Empowered by their exceptional capabilities to understand and generate human-like text, these models are being increasingly integrated into…

Cryptography and Security · Computer Science 2024-10-02 Zhiyuan Yu , Xiaogeng Liu , Shunning Liang , Zach Cameron , Chaowei Xiao , Ning Zhang

Text-to-image (T2I) models such as Stable Diffusion have advanced rapidly and are now widely used in content creation. However, these models can be misused to generate harmful content, including nudity or violence, posing significant safety…

Cryptography and Security · Computer Science 2025-06-13 Zilong Wang , Xiang Zheng , Xiaosen Wang , Bo Wang , Xingjun Ma , Yu-Gang Jiang

Text-to-Image(T2I) models typically deploy safety filters to prevent the generation of sensitive images. Unfortunately, recent jailbreaking attack methods manually design instructions for the LLM to generate adversarial prompts, which…

Cryptography and Security · Computer Science 2025-11-24 Chenyu Zhang , Lanjun Wang , Yiwen Ma , Wenhui Li , An-An Liu

Text-to-image (T2I) generative models have revolutionized content creation by transforming textual descriptions into high-quality images. However, these models are vulnerable to jailbreaking attacks, where carefully crafted prompts bypass…

Cryptography and Security · Computer Science 2025-06-26 Yingkai Dong , Xiangtao Meng , Ning Yu , Zheng Li , Shanqing Guo

Large Language models (LLMs) are transforming digital usage, particularly in text generation, image creation, information retrieval and code development. ChatGPT, launched by OpenAI in November 2022, quickly became a reference, prompting…

Cryptography and Security · Computer Science 2025-06-13 Rafaël Nouailles

Jailbreak attacks on multimodal AI systems remain underexplored, even though unsafe image generation can have more severe consequences than unsafe text and current defenses are relatively immature. We introduce PAST2HARM, a simple yet…

Computation and Language · Computer Science 2026-05-28 Snehasis Mukhopadhyay

Diffusion models have recently achieved remarkable advancements in terms of image quality and fidelity to textual prompts. Concurrently, the safety of such generative models has become an area of growing concern. This work introduces a…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Tong Liu , Zhixin Lai , Jiawen Wang , Gengyuan Zhang , Shuo Chen , Philip Torr , Vera Demberg , Volker Tresp , Jindong Gu

Text-to-image generative models are widely deployed in creative tools and online platforms. To mitigate misuse, these systems rely on safety filters and moderation pipelines that aim to block harmful or policy violating content. In this…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Ahmed B Mustafa , Zihan Ye , Yang Lu , Michael P Pound , Shreyank N Gowda

In recent years, Text-to-Image (T2I) models have garnered significant attention due to their remarkable advancements. However, security concerns have emerged due to their potential to generate inappropriate or Not-Safe-For-Work (NSFW)…

Computer Vision and Pattern Recognition · Computer Science 2025-02-11 Yihao Huang , Le Liang , Tianlin Li , Xiaojun Jia , Run Wang , Weikai Miao , Geguang Pu , Yang Liu

Modern text-to-image (T2I) models can now render legible, paragraph-length text, enabling a fundamentally new class of misuse. We identify and formalize the inscriptive jailbreak, where an adversary coerces a T2I system into generating…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Zonghao Ying , Haowen Dai , Lianyu Hu , Zonglei Jing , Quanchen Zou , Yaodong Yang , Aishan Liu , Xianglong Liu

Modern text-to-image (T2I) generation systems (e.g., DALL$\cdot$E 3) exploit the memory mechanism, which captures key information in multi-turn interactions for faithful generation. Despite its practicality, the security analyses of this…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Shiqian Zhao , Jiayang Liu , Yiming Li , Runyi Hu , Xiaojun Jia , Wenshu Fan , Xiao Bao , Xinfeng Li , Jie Zhang , Wei Dong , Tianwei Zhang , Luu Anh Tuan

Large Language Models (LLMs) have revolutionized Artificial Intelligence (AI) services due to their exceptional proficiency in understanding and generating human-like text. LLM chatbots, in particular, have seen widespread adoption,…

Cryptography and Security · Computer Science 2024-02-14 Gelei Deng , Yi Liu , Yuekang Li , Kailong Wang , Ying Zhang , Zefeng Li , Haoyu Wang , Tianwei Zhang , Yang Liu

Large Language Models (LLMs), such as ChatGPT, encounter `jailbreak' challenges, wherein safeguards are circumvented to generate ethically harmful prompts. This study introduces a straightforward black-box method for efficiently crafting…

Computation and Language · Computer Science 2024-04-25 Kazuhiro Takemoto

Text-to-image (T2I) models can be maliciously used to generate harmful content such as sexually explicit, unfaithful, and misleading or Not-Safe-for-Work (NSFW) images. Previous attacks largely depend on the availability of the diffusion…

Cryptography and Security · Computer Science 2025-05-27 Jiachen Ma , Yijiang Li , Zhiqing Xiao , Anda Cao , Jie Zhang , Chao Ye , Junbo Zhao

Text-to-Image models may generate harmful content, such as pornographic images, particularly when unsafe prompts are submitted. To address this issue, safety filters are often added on top of text-to-image models, or the models themselves…

Cryptography and Security · Computer Science 2026-01-09 Zhengyuan Jiang , Yuepeng Hu , Yuchen Yang , Yinzhi Cao , Neil Zhenqiang Gong

Text-to-image (T2I) models commonly incorporate defense mechanisms to prevent the generation of sensitive images. Unfortunately, recent jailbreak attacks have shown that adversarial prompts can effectively bypass these mechanisms and induce…

Cryptography and Security · Computer Science 2026-03-25 Chenyu Zhang , Lanjun Wang , Yiwen Ma , Wenhui Li , Yi Tu , An-An Liu

The rapid development of generative artificial intelligence has made text to video models essential for building future multimodal world simulators. However, these models remain vulnerable to jailbreak attacks, where specially crafted…

Cryptography and Security · Computer Science 2025-04-29 Siyuan Liang , Jiayang Liu , Jiecheng Zhai , Tianmeng Fang , Rongcheng Tu , Aishan Liu , Xiaochun Cao , Dacheng Tao

In recent years, fueled by the rapid advancement of diffusion models, text-to-video (T2V) generation models have achieved remarkable progress, with notable examples including Pika, Luma, Kling, and Open-Sora. Although these models exhibit…

Computer Vision and Pattern Recognition · Computer Science 2025-06-18 Jiayang Liu , Siyuan Liang , Shiqian Zhao , Rongcheng Tu , Wenbo Zhou , Aishan Liu , Dacheng Tao , Siew Kei Lam
‹ Prev 1 2 3 10 Next ›