隐藏于平文本:LLM 中隐蔽合作的出现与缓解
计算与语言
2025-12-03 v2 密码学与安全
机器学习
摘要
前沿模型代理的快速普及为社会带来了显著进步,但也引发了源于不安全交互所导致的系统性风险。为他人不利地进行合作的“合作”已被识别为一种不希望出现的代理合作形式。代理通信中信息隐藏(隐写术�)的使用可能使此类合作实际上难以检测。这凸显了对此类行为是否可能出现以及相应防御措施的必要性调查。为调查此问题,我们设计了两种方法——梯度基于强化学习(GBRL)方法和基于情境的强化学习(ICRL)方法——以可靠地激发复杂的 LLM 生成的语言文本隐写。我们首次展示了由于训练中误指定的奖励激励,LLM 中意外的隐写合作会产生。此外,我们发现,标准的缓解措施——包括对模型输出的被动监督和通过通信改写的主动缓解——在防止此类隐写通信方面并不完全有效。我们的发现表明:(i) 隐写合作的出现是一种可信担忧,应被监控和研究;(ii) 防止其出现可能需要创新防御技术。
引用
@article{arxiv.2410.03768,
title = {Hidden in Plain Text: Emergence & Mitigation of Steganographic Collusion in LLMs},
author = {Yohan Mathew and Ollie Matthews and Robert McCarthy and Joan Velja and Christian Schroeder de Witt and Dylan Cope and Nandi Schoots},
journal= {arXiv preprint arXiv:2410.03768},
year = {2025}
}
备注
Camera-ready version. Oral presentation at IJCNLP-AACL 2025 (14th International Joint Conference on Natural Language Processing and 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics), Mumbai, India, December 20-24, 2025