中文

Legilimens:面向大语言模型服务的实用统一内容审查

计算与语言 2024-09-06 v2

摘要

鉴于大语言模型 (LLM) 生成的unsafe content 的社会影响,确保 LLM 服务符合安全标准是 LLM 服务提供商的关键关注点。常见的内容审查方法受限于有效性与效率之间的困境,其中简单模型脆弱而高级模型消耗大量计算资源。本文首次揭示,尽管初始微调目标是对话而非内容审查,从聊天导向的 LLM 中提取概念特征即可实现有效且高效的内容审查。我们提出一种实用且统一的 LLM 服务内容审查框架,命名为 Legilimens,具备有效性和效率。我们的 red-team model-based 数据增强提高了 Legilimens 针对 state-of-the-art 恶作剧的鲁棒性。此外,我们开发了一个框架来理论分析 Legilimens 相较于其他方法的成本效益。我们在五个主流 LLM 上、十七个数据集和九种恶作剧方法上进行了广泛实验,以验证 Legilimens 针对 normal 和 adaptive 对手的有效性、效率和鲁棒性。与商业和学术基准的比较表明,Legilimens 性能卓越。此外,我们确认 Legilimens 可应用于 few-shot 场景并扩展到 multi-label 分类任务。

关键词

引用

@article{arxiv.2408.15488,
  title  = {Legilimens: Practical and Unified Content Moderation for Large Language Model Services},
  author = {Jialin Wu and Jiangyi Deng and Shengyuan Pang and Yanjiao Chen and Jiayang Xu and Xinfeng Li and Wenyuan Xu},
  journal= {arXiv preprint arXiv:2408.15488},
  year   = {2024}
}

备注

Accepted by ACM Conference on Computer and Communications Security (CCS) 2024