知识返回定向提示 (KROP)
密码学与安全
2024-06-19 v1 机器学习
摘要
当今许多大型语言模型 (LLM) 及其驱动的应用程序使用的某种形式的提示过滤或对齐措施旨在保护其完整性,但这些措施并非万无一失。本文引入 KROP,一种能够混淆提示注入攻击的提示注入技术,使其对大多数安全措施几乎不可检测。
引用
@article{arxiv.2406.11880,
title = {Knowledge Return Oriented Prompting (KROP)},
author = {Jason Martin and Kenneth Yeung},
journal= {arXiv preprint arXiv:2406.11880},
year = {2024}
}