忽略此标题与HackAPrompt:通过全球规模提示黑客竞赛揭示大语言模型的系统性漏洞
密码学与安全
2024-03-05 v3 人工智能
计算与语言
摘要
大语言模型(LLMs)被部署在具有直接用户交互的交互式场景中,例如聊天机器人和写作助手。这些部署易受提示注入和越狱(统称为提示黑客)的攻击,即模型被操纵以忽略其原始指令并遵循潜在的恶意指令。尽管被广泛认为是一个重大的安全威胁,但目前缺乏关于提示黑客的大规模资源和定量研究。为弥补这一空白,我们发起了一场全球提示黑客竞赛,允许自由形式的人类输入攻击。我们针对三个最先进的LLMs收集了60万以上的对抗提示。我们描述了该数据集,其经验性地验证了当前LLMs确实可通过提示黑客被操纵。我们还提出了对抗提示类型的全面分类本体。
引用
@article{arxiv.2311.16119,
title = {Ignore This Title and HackAPrompt: Exposing Systemic Vulnerabilities of LLMs through a Global Scale Prompt Hacking Competition},
author = {Sander Schulhoff and Jeremy Pinto and Anaum Khan and Louis-François Bouchard and Chenglei Si and Svetlina Anati and Valen Tagliabue and Anson Liu Kost and Christopher Carnahan and Jordan Boyd-Graber},
journal= {arXiv preprint arXiv:2311.16119},
year = {2024}
}
备注
34 pages, 8 figures Codebase: https://github.com/PromptLabs/hackaprompt Dataset: https://huggingface.co/datasets/hackaprompt/hackaprompt-dataset/blob/main/README.md Playground: https://huggingface.co/spaces/hackaprompt/playground