Fact2Fiction:针对具智能体功能的事实核查系统的定向投毒攻击
密码学与安全
2025-11-18 v2 计算与语言
摘要
面向事实核查系统的攻击安全性至关重要,因为受损的核查系统可能放大误导信息,但目前研究 largely 未知。为填补这一差距,本工作提出一种针对此类事实核查系统的新型威胁模型,并介绍了首个针对 SOTA agentic fact-checking 系统的投毒攻击框架 Fact2Fiction。Fact2Fiction 利用大语言模型模拟分解策略,并利用系统生成的 justifications 来构建针对性的恶意证据,以破坏子论点核查。大量实验表明,Fact2Fiction 在各种投毒预算下,相较于 SOTA 攻击实现了 8.9%--21.2% 的更高攻击成功率,揭示了现有事实核查系统的安全漏洞,凸显了需要防御措施的必要性。
引用
@article{arxiv.2508.06059,
title = {Fact2Fiction: Targeted Poisoning Attack to Agentic Fact-checking System},
author = {Haorui He and Yupeng Li and Bin Benjamin Zhu and Dacheng Wen and Reynold Cheng and Francis C. M. Lau},
journal= {arXiv preprint arXiv:2508.06059},
year = {2025}
}
备注
Accepted by AAAI 2026 (Oral). Code available at: https://trustworthycomp.github.io/Fact2Fiction/