Cryptography and Security · Computer Science
TrojLLM: A Black-box Trojan Prompt Attack on Large Language Models
Jiaqi Xue, Mengxin Zheng, Ting Hua, Yilin Shen +3
2023-11-01
Computers and Society · Computer Science
CodeGuard: Improving LLM Guardrails in CS Education
Nishat Raihan, Noah Erdachew, Jayoti Devi, Joanna C. S. Santos +1
2026-02-04
Cryptography and Security · Computer Science
Trojaning Language Models for Fun and Profit
Xinyang Zhang, Zheng Zhang, Shouling Ji, Ting Wang
2021-03-12
Computation and Language · Computer Science
Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks
Abhinav Rao, Sachin Vashistha, Atharva Naik, Somak Aditya +1
2024-03-28
Computation and Language · Computer Science
Does Safety Training of LLMs Generalize to Semantically Related Natural Prompts?
Sravanti Addepalli, Yerram Varun, Arun Suggala, Karthikeyan Shanmugam +1
2025-03-26
Cryptography and Security · Computer Science
Prompt Inject Detection with Generative Explanation as an Investigative Tool
Jonathan Pan, Swee Liang Wong, Yidi Yuan, Xin Wei Chia
2025-02-18
Computation and Language · Computer Science
GradSafe: Detecting Jailbreak Prompts for LLMs via Safety-Critical Gradient Analysis
Yueqi Xie, Minghong Fang, Renjie Pi, Neil Gong
2024-05-31
Cryptography and Security · Computer Science
Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs through a Global Scale Prompt Hacking Competition
Sander Schulhoff, Jeremy Pinto, Anaum Khan, Louis-François Bouchard +6
2024-03-05
Cryptography and Security · Computer Science
Prefix Guidance: A Steering Wheel for Large Language Models to Defend Against Jailbreak Attacks
Jiawei Zhao, Kejiang Chen, Xiaojian Yuan, Weiming Zhang
2024-08-23
Computation and Language · Computer Science
UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models
Huawei Lin, Yingjie Lao, Tong Geng, Tan Yu +1
2025-02-19
Computer Vision and Pattern Recognition · Computer Science
PromptGuard: An Orchestrated Prompting Framework for Principled Synthetic Text Generation for Vulnerable Populations using LLMs with Enhanced Safety, Fairness, and Controllability
Tung Vu, Lam Nguyen, Quynh Dao
2025-09-12
Computation and Language · Computer Science
Certifying LLM Safety against Adversarial Prompting
Aounon Kumar, Chirag Agarwal, Suraj Srinivas, Aaron Jiaxun Li +2
2025-02-06
Computation and Language · Computer Science
LLMs are Vulnerable to Malicious Prompts Disguised as Scientific Language
Yubin Ge, Neeraja Kirtane, Hao Peng, Dilek Hakkani-Tür
2025-02-19
Software Engineering · Computer Science
Leveraging Large Language Models for Command Injection Vulnerability Analysis in Python: An Empirical Study on Popular Open-Source Projects
Yuxuan Wang, Jingshu Chen, Qingyang Wang
2025-05-22
Artificial Intelligence · Computer Science
TroubleLLM: Align to Red Team Expert
Zhuoer Xu, Jianping Zhang, Shiwen Cui, Changhua Meng +1
2024-03-05
Computation and Language · Computer Science
Multilingual Jailbreak Challenges in Large Language Models
Yue Deng, Wenxuan Zhang, Sinno Jialin Pan, Lidong Bing
2024-03-05
Computation and Language · Computer Science
Do Prompts Guarantee Safety? Mitigating Toxicity from LLM Generations through Subspace Intervention
Himanshu Singh, Ziwei Xu, A. V. Subramanyam, Mohan Kankanhalli
2026-02-09
Computation and Language · Computer Science
PromptAttack: Prompt-based Attack for Language Models via Gradient Search
Yundi Shi, Piji Li, Changchun Yin, Zhaoyang Han +2
2022-09-07
Machine Learning · Computer Science
Are My Optimized Prompts Compromised? Exploring Vulnerabilities of LLM-based Optimizers
Andrew Zhao, Reshmi Ghosh, Vitor Carvalho, Emily Lawton +3
2026-01-14
Computation and Language · Computer Science
LLM Self Defense: By Self Examination, LLMs Know They Are Being Tricked
Mansi Phute, Alec Helbling, Matthew Hull, ShengYun Peng +3
2024-05-03
Cryptography and Security · Computer Science
Misleading Large Language Models used (or misused) in Scientific Peer-Reviewing via Hidden Prompt-Injection Attacks
Matteo Gioele Collu, Umberto Salviati, Roberto Confalonieri, Mauro Conti +1
2026-03-31
Computation and Language · Computer Science
Trojan Detection in Large Language Models: Insights from The Trojan Detection Challenge
Narek Maloyan, Ekansh Verma, Bulat Nutfullin, Bislan Ashinov
2024-04-23