Cryptography and Security · Computer Science
SOS! Soft Prompt Attack Against Open-Source Large Language Models
Ziqing Yang, Michael Backes, Yang Zhang, Ahmed Salem
2024-07-04
Machine Learning · Computer Science
Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space
Leo Schwinn, David Dobre, Sophie Xhonneux, Gauthier Gidel +1
2025-04-17
Cryptography and Security · Computer Science
Is Your Prompt Safe? Investigating Prompt Injection Attacks Against Open-Source LLMs
Jiawen Wang, Pritha Gupta, Ivan Habernal, Eyke Hüllermeier
2025-05-21
Machine Learning · Computer Science
StructTransform: A Scalable Attack Surface for Safety-Aligned Large Language Models
Shehel Yoosuf, Temoor Ali, Ahmed Lekssays, Mashael AlSabah +1
2025-07-04
Machine Learning · Computer Science
Safe LoRA: the Silver Lining of Reducing Safety Risks when Fine-tuning Large Language Models
Chia-Yi Hsu, Yu-Lin Tsai, Chih-Hsun Lin, Pin-Yu Chen +2
2025-01-07
Computation and Language · Computer Science
LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
Mansi Phute, Alec Helbling, Matthew Hull, ShengYun Peng +3
2024-05-03
Computation and Language · Computer Science
Towards Understanding the Fragility of Multilingual LLMs against Fine-Tuning Attacks
Samuele Poppi, Zheng-Xin Yong, Yifei He, Bobbie Chern +3
2025-03-03
Machine Learning · Computer Science
Rethinking Safety in LLM Fine-tuning: An Optimization Perspective
Minseon Kim, Jin Myung Kwak, Lama Alssum, Bernard Ghanem +4
2025-08-19
Computation and Language · Computer Science
CodeAttack: Revealing Safety Generalization Challenges of Large Language Models via Code Completion
Qibing Ren, Chang Gao, Jing Shao, Junchi Yan +3
2024-09-17
Machine Learning · Computer Science
Does Low Rank Adaptation Lead to Lower Robustness against Training-Time Attacks?
Zi Liang, Haibo Hu, Qingqing Ye, Yaxin Xiao +1
2025-05-20
Computation and Language · Computer Science
Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen +3
2023-10-06
Cryptography and Security · Computer Science
Harnessing Task Overload for Scalable Jailbreak Attacks on Large Language Models
Yiting Dong, Guobin Shen, Dongcheng Zhao, Xiang He +1
2024-10-08
Cryptography and Security · Computer Science
Meta SecAlign: A Secure Foundation LLM Against Prompt Injection Attacks
Sizhe Chen, Arman Zharmagambetov, David Wagner, Chuan Guo
2026-02-09
Cryptography and Security · Computer Science
Jailbreaking Attacks vs. Content Safety Filters: How Far Are We in the LLM Safety Arms Race?
Yuan Xin, Dingfan Chen, Linyi Yang, Michael Backes +1
2026-01-01
Software Engineering · Computer Science
From Rookie to Expert: Manipulating LLMs for Automated Vulnerability Exploitation in Enterprise Software
Moustapha Awwalou Diouf, Maimouna Tamah Diao, Iyiola Emmanuel Olatunji, Abdoul Kader Kaboré +5
2026-04-24
Cryptography and Security · Computer Science
Defense Against Prompt Injection Attack by Leveraging Attack Techniques
Yulin Chen, Haoran Li, Zihao Zheng, Yangqiu Song +2
2025-08-05
Cryptography and Security · Computer Science
Safety Alignment Should Be Made More Than Just a Few Tokens Deep
Xiangyu Qi, Ashwinee Panda, Kaifeng Lyu, Xiao Ma +4
2024-06-11
Cryptography and Security · Computer Science
Death by a Thousand Prompts: Open Model Vulnerability Analysis
Amy Chang, Nicholas Conley, Harish Santhanalakshmi Ganesan, Adam Swanda
2025-11-06
Cryptography and Security · Computer Science
Beyond the Benchmark: Innovative Defenses Against Prompt Injection Attacks
Safwan Shaheer, G. M. Refatul Islam, Mohammad Rafid Hamid, Tahsin Zaman Jilan
2025-12-19
Computation and Language · Computer Science
PromptAttack: Prompt-based Attack for Language Models via Gradient Search
Yundi Shi, Piji Li, Changchun Yin, Zhaoyang Han +2
2022-09-07