Artificial Intelligence · Computer Science
The Struggle Between Continuation and Refusal: A Mechanistic Analysis of the Continuation-Triggered Jailbreak in LLMs
Yonghong Deng, Zhen Yang, Ping Jian, Xinyue Zhang +2
2026-03-10
Cryptography and Security · Computer Science
The VLLM Safety Paradox: Dual Ease in Jailbreak Attack and Defense
Yangyang Guo, Fangkai Jiao, Liqiang Nie, Mohan Kankanhalli
2025-03-07
Cryptography and Security · Computer Science
JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
Junjie Chu, Yugeng Liu, Ziqing Yang, Xinyue Shen +2
2025-05-27
Cryptography and Security · Computer Science
SelfDefend: LLMs Can Defend Themselves against Jailbreaking in a Practical Manner
Xunguang Wang, Daoyuan Wu, Zhenlan Ji, Zongjie Li +6
2025-02-06
Computation and Language · Computer Science
A Domain-Based Taxonomy of Jailbreak Vulnerabilities in Large Language Models
Carlos Peláez-González, Andrés Herrera-Poyatos, Cristina Zuheros, David Herrera-Poyatos +2
2025-04-08
Computation and Language · Computer Science
Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks
Abhinav Rao, Sachin Vashistha, Atharva Naik, Somak Aditya +1
2024-03-28
Cryptography and Security · Computer Science
Jailbreaking Attacks vs. Content Safety Filters: How Far Are We in the LLM Safety Arms Race?
Yuan Xin, Dingfan Chen, Linyi Yang, Michael Backes +1
2026-01-01
Computation and Language · Computer Science
Diversity Helps Jailbreak Large Language Models
Weiliang Zhao, Daniel Ben-Levi, Wei Hao, Junfeng Yang +1
2025-05-13
Computation and Language · Computer Science
EasyJailbreak: A Unified Framework for Jailbreaking Large Language Models
Weikang Zhou, Xiao Wang, Limao Xiong, Han Xia +17
2024-03-20
Machine Learning · Computer Science
Jailbreak Defense in a Narrow Domain: Limitations of Existing Methods and a New Transcript-Classifier Approach
Tony T. Wang, John Hughes, Henry Sleight, Rylan Schaeffer +6
2024-12-04
Computation and Language · Computer Science
Machine Learning for Detection and Analysis of Novel LLM Jailbreaks
John Hawkins, Aditya Pramar, Rodney Beard, Rohitash Chandra
2025-10-13
Cryptography and Security · Computer Science
Emerging Vulnerabilities in Frontier Models: Multi-Turn Jailbreak Attacks
Tom Gibbs, Ethan Kosak-Hine, George Ingebretsen, Jason Zhang +5
2024-09-04
Cryptography and Security · Computer Science
What Features in Prompts Jailbreak LLMs? Investigating the Mechanisms Behind Attacks
Nathalie Kirch, Constantin Weisser, Severin Field, Helen Yannakoudakis +1
2025-11-04
Cryptography and Security · Computer Science
Confusion is the Final Barrier: Rethinking Jailbreak Evaluation and Investigating the Real Misuse Threat of LLMs
Yu Yan, Sheng Sun, Zhe Wang, Yijun Lin +5
2025-09-16
Cryptography and Security · Computer Science
You Can't Eat Your Cake and Have It Too: The Performance Degradation of LLMs with Jailbreak Defense
Wuyuao Mai, Geng Hong, Pei Chen, Xudong Pan +4
2025-01-22
Cryptography and Security · Computer Science
JailPO: A Novel Black-box Jailbreak Framework via Preference Optimization against Aligned LLMs
Hongyi Li, Jiawei Ye, Jie Wu, Tianjie Yan +2
2024-12-23
Computation and Language · Computer Science
Dark LLMs: The Growing Threat of Unaligned AI Models
Michael Fire, Yitzhak Elbazis, Adi Wasenstein, Lior Rokach
2025-05-16
Cryptography and Security · Computer Science
JailbreakLens: Interpreting Jailbreak Mechanism in the Lens of Representation and Circuit
Zeqing He, Zhibo Wang, Zhixuan Chu, Huiyu Xu +3
2025-04-25
Cryptography and Security · Computer Science
How Jailbreak Defenses Work and Ensemble? A Mechanistic Investigation
Zhuohang Long, Siyuan Wang, Shujun Liu, Yuhang Lai +2
2025-02-21
Computation and Language · Computer Science
Uncovering the Persuasive Fingerprint of LLMs in Jailbreaking Attacks
Havva Alizadeh Noughabi, Julien Serbanescu, Fattane Zarrinkalam, Ali Dehghantanha
2025-10-28