Machine Learning · Computer Science
Attacking Large Language Models with Projected Gradient Descent
Simon Geisler, Tom Wollschläger, M. H. I. Abdalla, Johannes Gasteiger +1
2025-03-04
Computation and Language · Computer Science
Bridge the Gap Between CV and NLP! A Gradient-based Textual Adversarial Attack Framework
Lifan Yuan, Yichi Zhang, Yangyi Chen, Wei Wei
2023-06-09
Computer Vision and Pattern Recognition · Computer Science
Enhancing Targeted Adversarial Attacks on Large Vision-Language Models via Intermediate Projector
Yiming Cao, Yanjie Li, Kaisheng Liang, Bin Xiao
2025-09-25
Computation and Language · Computer Science
Palisade -- Prompt Injection Detection Framework
Sahasra Kokkula, Somanathan R, Nandavardhan R, Aashishkumar +1
2024-10-29
Cryptography and Security · Computer Science
Super Suffixes: Bypassing Text Generation Alignment and Guard Models Simultaneously
Andrew Adiletta, Kathryn Adiletta, Kemal Derya, Berk Sunar
2025-12-15
Machine Learning · Computer Science
Universal and Transferable Adversarial Attack on Large Language Models Using Exponentiated Gradient Descent
Sajib Biswas, Mao Nishino, Samuel Jacob Chacko, Xiuwen Liu
2025-08-21
Machine Learning · Computer Science
Adversarial Attack on Large Language Models using Exponentiated Gradient Descent
Sajib Biswas, Mao Nishino, Samuel Jacob Chacko, Xiuwen Liu
2025-05-16
Computation and Language · Computer Science
Exploiting the Index Gradients for Optimization-Based Jailbreaking on Large Language Models
Jiahui Li, Yongchang Hao, Haoyu Xu, Xing Wang +1
2024-12-17
Computer Vision and Pattern Recognition · Computer Science
Mitigating Adversarial Attacks in LLMs through Defensive Suffix Generation
Minkyoung Kim, Yunha Kim, Hyeram Seo, Heejung Choi +8
2024-12-19
Computation and Language · Computer Science
ASETF: A Novel Method for Jailbreak Attack on LLMs through Translate Suffix Embeddings
Hao Wang, Hao Li, Minlie Huang, Lei Sha
2024-06-05
Computation and Language · Computer Science
Multi-Objective Large Language Model Unlearning
Zibin Pan, Shuwen Zhang, Yuesheng Zheng, Chi Li +2
2025-01-07
Machine Learning · Computer Science
Hijacking Large Language Models via Adversarial In-Context Learning
Xiangyu Zhou, Yao Qiang, Saleh Zare Zade, Prashant Khanduri +1
2025-05-30
Cryptography and Security · Computer Science
TopicAttack: An Indirect Prompt Injection Attack via Topic Transition
Yulin Chen, Haoran Li, Yuexin Li, Yue Liu +2
2025-10-07
Artificial Intelligence · Computer Science
MAPGD: Multi-Agent Prompt Gradient Descent for Collaborative Prompt Optimization
Yichen Han, Yuhang Han, Siteng Huang, Guanyu Liu +6
2026-02-04
Cryptography and Security · Computer Science
Misleading Large Language Models used (or misused) in Scientific Peer-Reviewing via Hidden Prompt-Injection Attacks
Matteo Gioele Collu, Umberto Salviati, Roberto Confalonieri, Mauro Conti +1
2026-03-31
Information Retrieval · Computer Science
GRADA: Graph-based Reranking against Adversarial Documents Attack
Jingjie Zheng, Aryo Pradipta Gema, Giwon Hong, Xuanli He +3
2025-09-19
Artificial Intelligence · Computer Science
Automatic and Universal Prompt Injection Attacks against Large Language Models
Xiaogeng Liu, Zhiyuan Yu, Yizhe Zhang, Ning Zhang +1
2024-03-11
Computer Vision and Pattern Recognition · Computer Science
Adversarial Attacks Against MLLMs via Progressive Resolution Processing and Adaptive Feature Alignment
Haobo Wang, Xiaorong Ma, Weiqi Luo, Xiaojun Jia +1
2026-05-12
Cryptography and Security · Computer Science
PLA: Prompt Learning Attack against Text-to-Image Generative Models
Xinqi Lyu, Yihao Liu, Yanjie Li, Bin Xiao
2025-08-07