在多模态 LLM 中滥用图像与声音进行间接指令注入
密码学与安全
2023-10-04 v4 人工智能
计算与语言
机器学习
摘要
我们展示了图像和声音如何用于多模态 LLM 中的间接提示与指令注入。攻击者生成对应于提示的对抗扰动并将其混合进图像或音频记录中。当用户向(未修改、良性的)模型询问被扰动的图像或音频时,该扰动引导模型输出攻击者选定的文本和/或使后续对话遵循攻击者的指令。我们通过对 LLaVa 和 PandaGPT 的若干概念验证示例说明了此攻击。
引用
@article{arxiv.2307.10490,
title = {Abusing Images and Sounds for Indirect Instruction Injection in Multi-Modal LLMs},
author = {Eugene Bagdasaryan and Tsung-Yin Hsieh and Ben Nassi and Vitaly Shmatikov},
journal= {arXiv preprint arXiv:2307.10490},
year = {2023}
}