中文

生成模型的对抗性域外样本

机器学习 2019-05-15 v2 机器学习

摘要

深度生成模型正迅速成为研究人员与开发者的常用工具。然而,正如对判别模型族已详尽展示的那样,深度神经网络在测试时的推断无法被完全控制,且错误行为可由攻击者诱发。在本工作中,我们展示恶意用户如何通过向预训练生成器输入合适的对抗性输入,迫使其复现任意数据实例。此外,我们表明这些对抗性隐向量可被塑造成在统计上与真实输入集不可区分。所提出的攻击技术针对各种采用不同架构、训练过程以及条件与非条件设置的 GAN 图像生成器进行了评估。

关键词

引用

@article{arxiv.1903.02926,
  title  = {Adversarial Out-domain Examples for Generative Models},
  author = {Dario Pasquini and Marco Mingione and Massimo Bernaschi},
  journal= {arXiv preprint arXiv:1903.02926},
  year   = {2019}
}

备注

accepted in proceedings of the Workshop on Machine Learning for Cyber-Crime Investigation and Cybersecurity