攻击遇见可解释性(AmI)的评估与发现
密码学与安全
2026-04-14 v4
摘要
为研究模型解释在检测对抗样本中的有效性,我们复现了两篇论文的结果:Attacks Meet Interpretability: Attribute-steered Detection of Adversarial Samples 与 Is AmI (Attacks Meet Interpretability) Robust to Adversarial Samples,并通过实验和案例研究指出了两项工作的局限性。我们发现攻击遇见可解释性(AmI)高度依赖于超参数的选择。因此,在另一种超参数选择下,AmI仍能检测Nicholas Carlini的攻击。最后,我们就未来诸如AmI等防御技术评估的工作提出了建议。
引用
@article{arxiv.2310.08808,
title = {Attacks Meet Interpretability (AmI) Evaluation and Findings},
author = {Qian Ma and Ziping Ye and Shagufta Mehnaz},
journal= {arXiv preprint arXiv:2310.08808},
year = {2026}
}
备注
Experiments issues need to be fixed