中文

加速低查询场景下的针对性硬标签对抗攻击

计算机视觉与模式识别 2026-04-23 v3 机器学习

摘要

用于图像分类的深度神经网络仍然 vulnerable to 对抗样本——小幅、不可感知的扰动,可诱导误分类。在黑箱设置下,仅访问最终预测的情况下,制作旨在误分类到特定目标类别的针对性攻击尤其具有挑战性 due to 决策区域狭窄。当前最先进的方法常利用源图像与目标图像之间决策边界的几何性质,而未融合图像本身的信息。相反,我们提出了Targeted Edge-informed Attack (TEA),一种新型攻击方法,利用目标图像的边缘信息对其进行精确扰动,从而产生接近源图像且仍实现所需目标分类的对抗图像。我们的方法在不同模型上处于低查询设置下(使用查询次数几乎减少约70%)均显著优于当前最先进的方法,这在实际应用中尤其 relevant due to 查询有限和黑箱访问限制。此外,TEA高效生成合适的对抗样本,为 established geometry-based攻击提供改进的目标初始化。

关键词

引用

@article{arxiv.2505.16313,
  title  = {Accelerating Targeted Hard-Label Adversarial Attacks in Low-Query Black-Box Settings},
  author = {Arjhun Swaminathan and Mete Akgün},
  journal= {arXiv preprint arXiv:2505.16313},
  year   = {2026}
}

备注

This paper contains 10 pages, 8 figures and 8 tables. For associated supplementary code, see https://github.com/mdppml/TEA. This work has been accepted for publication at the IEEE Conference on Secure and Trustworthy Machine Learning (SaTML). The final version will be available on IEEE Xplore