Malla:揭秘现实世界中集成大型语言模型的恶意服务
密码学与安全
2024-08-21 v3 人工智能
摘要
地下领域利用大型语言模型(LLMs)提供恶意服务(即 Malla)的现象日益增多,加剧了网络威胁态势,并引发了关于 LLM 技术可信度的质疑。然而,目前鲜有研究从规模、影响和技术层面去理解这种新型网络犯罪。本文对 212 个现实世界中的 Malla 进行了首次系统性研究,揭示了它们在地下市场中的扩散情况及其运作模式。我们的研究披露了 Malla 生态系统,展现了其显著的增长及其对当今公共 LLM 服务的影响。通过对 212 个 Malla 的审查,我们发现了 Malla 使用的八个后端 LLM,以及 182 个用于规避公共 LLM API 防护措施的提示词(prompts)。我们进一步揭示了 Malla 采用的策略,包括滥用无审查 LLM 以及通过越狱提示词(jailbreak prompts)利用公共 LLM API。我们的发现有助于更好地理解网络犯罪分子对 LLM 的现实世界利用,并为对抗此类网络犯罪的策略提供了见解。
引用
@article{arxiv.2401.03315,
title = {Malla: Demystifying Real-world Large Language Model Integrated Malicious Services},
author = {Zilong Lin and Jian Cui and Xiaojing Liao and XiaoFeng Wang},
journal= {arXiv preprint arXiv:2401.03315},
year = {2024}
}
备注
Accepted at the 33rd USENIX Security Symposium (USENIX Security '24). The data and code are available at https://github.com/idllresearch/malicious-gpt