Cryptography and Security · Computer Science
Defeating Prompt Injections by Design
Edoardo Debenedetti, Ilia Shumailov, Tianqi Fan, Jamie Hayes +6
2025-06-25
Machine Learning · Computer Science
GuardVal: Dynamic Large Language Model Jailbreak Evaluation for Comprehensive Safety Testing
Peiyan Zhang, Haibo Jin, Liying Kang, Haohan Wang
2025-07-11
Cryptography and Security · Computer Science
Jailbreaking and Mitigation of Vulnerabilities in Large Language Models
Benji Peng, Hanxuan Chen, Keyu Chen, Qian Niu +11
2026-05-29
Robotics · Computer Science
Enhancing Reliability in LLM-Integrated Robotic Systems: A Unified Approach to Security and Safety
Wenxiao Zhang, Xiangrui Kong, Conan Dewitt, Thomas Bräunl +1
2025-09-03
Machine Learning · Computer Science
Hijacking Large Language Models via Adversarial In-Context Learning
Xiangyu Zhou, Yao Qiang, Saleh Zare Zade, Prashant Khanduri +1
2025-05-30
Computation and Language · Computer Science
AttackEval: How to Evaluate the Effectiveness of Jailbreak Attacking on Large Language Models
Dong Shu, Chong Zhang, Mingyu Jin, Zihao Zhou +2
2025-03-19
Cryptography and Security · Computer Science
Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs through a Global Scale Prompt Hacking Competition
Sander Schulhoff, Jeremy Pinto, Anaum Khan, Louis-François Bouchard +6
2024-03-05
Cryptography and Security · Computer Science
LLM-enabled Applications Require System-Level Threat Monitoring
Yedi Zhang, Haoyu Wang, Xianglin Yang, Jin Song Dong +1
2026-02-24
Cryptography and Security · Computer Science
Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models
Yakai Li, Jiekang Hu, Weiduan Sang, Luping Ma +6
2025-08-27
Computation and Language · Computer Science
Certifying LLM Safety against Adversarial Prompting
Aounon Kumar, Chirag Agarwal, Suraj Srinivas, Aaron Jiaxun Li +2
2025-02-06
Cryptography and Security · Computer Science
Towards Reliable and Practical LLM Security Evaluations via Bayesian Modelling
Mary Llewellyn, Annie Gray, Josh Collyer, Michael Harries
2025-10-08
Cryptography and Security · Computer Science
SoK: Prompt Hacking of Large Language Models
Baha Rababah, Shang, Wu, Matthew Kwiatkowski +2
2024-10-21
Cryptography and Security · Computer Science
When Prompts Become Payloads: A Framework for Mitigating SQL Injection Attacks in Large Language Model-Driven Applications
Farzad Nourmohammadzadeh Motlagh, Mehrdad Hajizadeh, Mehryar Majd, Pejman Najafi +2
2026-05-12
Computation and Language · Computer Science
Machine Learning for Detection and Analysis of Novel LLM Jailbreaks
John Hawkins, Aditya Pramar, Rodney Beard, Rohitash Chandra
2025-10-13
Cryptography and Security · Computer Science
Harnessing Task Overload for Scalable Jailbreak Attacks on Large Language Models
Yiting Dong, Guobin Shen, Dongcheng Zhao, Xiang He +1
2024-10-08
Cryptography and Security · Computer Science
Quantifying Frontier LLM Capabilities for Container Sandbox Escape
Rahul Marchand, Art O Cathain, Jerome Wynne, Philippos Maximos Giavridis +4
2026-03-04
Software Engineering · Computer Science
DLAP: A Deep Learning Augmented Large Language Model Prompting Framework for Software Vulnerability Detection
Yanjing Yang, Xin Zhou, Runfeng Mao, Jinwei Xu +4
2024-05-03
Cryptography and Security · Computer Science
MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots
Gelei Deng, Yi Liu, Yuekang Li, Kailong Wang +5
2024-02-14
Cryptography and Security · Computer Science
Analysis of LLMs Against Prompt Injection and Jailbreak Attacks
Piyush Jaiswal, Aaditya Pratap, Shreyansh Saraswati, Harsh Kasyap +1
2026-02-27
Machine Learning · Computer Science
SafeLLM: Unlearning Harmful Outputs from Large Language Models against Jailbreak Attacks
Xiangman Li, Xiaodong Wu, Qi Li, Jianbing Ni +1
2025-08-22
Cryptography and Security · Computer Science
LLMs Can Defend Themselves Against Jailbreaking in a Practical Manner: A Vision Paper
Daoyuan Wu, Shuai Wang, Yang Liu, Ning Liu
2024-03-05