English

LAMP: Learning Universal Adversarial Perturbations for Multi-Image Tasks via Pre-trained Models

Computer Vision and Pattern Recognition 2026-01-30 v1

Abstract

Multimodal Large Language Models (MLLMs) have achieved remarkable performance across vision-language tasks. Recent advancements allow these models to process multiple images as inputs. However, the vulnerabilities of multi-image MLLMs remain unexplored. Existing adversarial attacks focus on single-image settings and often assume a white-box threat model, which is impractical in many real-world scenarios. This paper introduces LAMP, a black-box method for learning Universal Adversarial Perturbations (UAPs) targeting multi-image MLLMs. LAMP applies an attention-based constraint that prevents the model from effectively aggregating information across images. LAMP also introduces a novel cross-image contagious constraint that forces perturbed tokens to influence clean tokens, spreading adversarial effects without requiring all inputs to be modified. Additionally, an index-attention suppression loss enables a robust position-invariant attack. Experimental results show that LAMP outperforms SOTA baselines and achieves the highest attack success rates across multiple vision-language tasks and models.

Keywords

Cite

@article{arxiv.2601.21220,
  title  = {LAMP: Learning Universal Adversarial Perturbations for Multi-Image Tasks via Pre-trained Models},
  author = {Alvi Md Ishmam and Najibul Haque Sarker and Zaber Ibn Abdul Hakim and Chris Thomas},
  journal= {arXiv preprint arXiv:2601.21220},
  year   = {2026}
}

Comments

Accepted in main technical track AAAI 2026

R2 v1 2026-07-01T09:24:56.524Z