中文

蓝队对抗函数调用智能体

密码学与安全 2026-01-15 v1 人工智能

摘要

我们提出了一项实验评估,测试了四个声称具备函数调用能力的开源 LLM 针对三种不同攻击的鲁棒性,并衡量了八种不同防御措施的有效性。结果表明,这些模型默认并不安全,且这些防御措施目前尚无法应用于真实场景。

关键词

引用

@article{arxiv.2601.09292,
  title  = {Blue Teaming Function-Calling Agents},
  author = {Greta Dolcetti and Giulio Zizzo and Sergio Maffeis},
  journal= {arXiv preprint arXiv:2601.09292},
  year   = {2026}
}

备注

This work has been accepted to appear at the AAAI 2026 Workshop on Trust and Control in Agentic AI (TrustAgent)