English

Scalable Backdoor Detection in Neural Networks

Computer Vision and Pattern Recognition 2020-06-11 v1

Abstract

Recently, it has been shown that deep learning models are vulnerable to Trojan attacks, where an attacker can install a backdoor during training time to make the resultant model misidentify samples contaminated with a small trigger patch. Current backdoor detection methods fail to achieve good detection performance and are computationally expensive. In this paper, we propose a novel trigger reverse-engineering based approach whose computational complexity does not scale with the number of labels, and is based on a measure that is both interpretable and universal across different network and patch types. In experiments, we observe that our method achieves a perfect score in separating Trojaned models from pure models, which is an improvement over the current state-of-the art method.

Keywords

Cite

@article{arxiv.2006.05646,
  title  = {Scalable Backdoor Detection in Neural Networks},
  author = {Haripriya Harikumar and Vuong Le and Santu Rana and Sourangshu Bhattacharya and Sunil Gupta and Svetha Venkatesh},
  journal= {arXiv preprint arXiv:2006.05646},
  year   = {2020}
}
R2 v1 2026-06-23T16:11:54.778Z