Cryptography and Security · Computer Science
Jailbreak Attacks and Defenses Against Large Language Models: A Survey
Sibo Yi, Yule Liu, Zhen Sun, Tianshuo Cong +4
2024-09-02
Cryptography and Security · Computer Science
Recent Advances in Attack and Defense Approaches of Large Language Models
Jing Cui, Yishi Xu, Zhewei Huang, Shuchang Zhou +2
2024-12-03
Cryptography and Security · Computer Science
Attack and defense techniques in large language models: A survey and new perspectives
Zhiyu Liao, Kang Chen, Yuanguo Lin, Kangkang Li +4
2025-05-05
Computation and Language · Computer Science
Building Safe GenAI Applications: An End-to-End Overview of Red Teaming for Large Language Models
Alberto Purpura, Sahil Wadhwa, Jesse Zymet, Akshay Gupta +4
2025-03-06
Cryptography and Security · Computer Science
Jailbreaking and Mitigation of Vulnerabilities in Large Language Models
Benji Peng, Hanxuan Chen, Keyu Chen, Qian Niu +11
2026-05-29
Computation and Language · Computer Science
Operationalizing a Threat Model for Red-Teaming Large Language Models (LLMs)
Apurv Verma, Satyapriya Krishna, Sebastian Gehrmann, Madhavan Seshadri +6
2025-12-30
Computation and Language · Computer Science
Revisiting Jailbreaking for Large Language Models: A Representation Engineering Perspective
Tianlong Li, Zhenghua Wang, Wenhao Liu, Muling Wu +5
2025-02-24
Computation and Language · Computer Science
Red Teaming Language Model Detectors with Language Models
Zhouxing Shi, Yihan Wang, Fan Yin, Xiangning Chen +2
2023-10-20
Artificial Intelligence · Computer Science
Defending Large Language Models Against Jailbreak Attacks via Layer-specific Editing
Wei Zhao, Zhe Li, Yige Li, Ye Zhang +1
2024-06-17
Cryptography and Security · Computer Science
Distract Large Language Models for Automatic Jailbreak Attack
Zeguan Xiao, Yan Yang, Guanhua Chen, Yun Chen
2024-10-01
Cryptography and Security · Computer Science
A Comprehensive Study of Jailbreak Attack versus Defense for Large Language Models
Zihao Xu, Yi Liu, Gelei Deng, Yuekang Li +1
2024-05-20
Computation and Language · Computer Science
From LLMs to MLLMs: Exploring the Landscape of Multimodal Jailbreaking
Siyuan Wang, Zhuohan Long, Zhihao Fan, Zhongyu Wei
2024-06-24
Cryptography and Security · Computer Science
Defending Large Language Models Against Jailbreak Exploits with Responsible AI Considerations
Ryan Wong, Hosea David Yu Fei Ng, Dhananjai Sharma, Glenn Jun Jie Ng +1
2025-11-25
Cryptography and Security · Computer Science
Breaking Down the Defenses: A Comparative Survey of Attacks on Large Language Models
Arijit Ghosh Chowdhury, Md Mofijul Islam, Vaibhav Kumar, Faysal Hossain Shezan +3
2024-03-26
Cryptography and Security · Computer Science
A Survey on Backdoor Threats in Large Language Models (LLMs): Attacks, Defenses, and Evaluations
Yihe Zhou, Tao Ni, Wei-Bin Lee, Qingchuan Zhao
2025-02-11
Cryptography and Security · Computer Science
Securing Large Language Models: Addressing Bias, Misinformation, and Prompt Attacks
Benji Peng, Keyu Chen, Ming Li, Pohsun Feng +4
2025-11-26
Computation and Language · Computer Science
Survey of Vulnerabilities in Large Language Models Revealed by Adversarial Attacks
Erfan Shayegani, Md Abdullah Al Mamun, Yu Fu, Pedram Zaree +2
2023-10-18
Cryptography and Security · Computer Science
Large Language Models in Cybersecurity: State-of-the-Art
Farzad Nourmohammadzadeh Motlagh, Mehrdad Hajizadeh, Mehryar Majd, Pejman Najafi +2
2024-02-05
Cryptography and Security · Computer Science
From LLMs to MLLMs to Agents: A Survey of Emerging Paradigms in Jailbreak Attacks and Defenses within LLM Ecosystem
Yanxu Mao, Tiehan Cui, Peipei Liu, Datao You +1
2025-08-04
Cryptography and Security · Computer Science
Securing Large Language Models: Threats, Vulnerabilities and Responsible Practices
Sara Abdali, Richard Anarfi, CJ Barberan, Jia He +1
2025-06-13
Artificial Intelligence · Computer Science
Defending Large Language Models Against Jailbreak Attacks via In-Decoding Safety-Awareness Probing
Yinzhi Zhao, Ming Wang, Shi Feng, Xiaocui Yang +2
2026-02-02
Cryptography and Security · Computer Science
A Red Teaming Roadmap Towards System-Level Safety
Zifan Wang, Christina Q. Knight, Jeremy Kritz, Willow E. Primack +1
2025-06-10