English

Prompting the Unseen: Detecting Hidden Backdoors in Black-Box Models

Computer Vision and Pattern Recognition 2025-04-08 v2 Artificial Intelligence Machine Learning

Abstract

Visual prompting (VP) is a new technique that adapts well-trained frozen models for source domain tasks to target domain tasks. This study examines VP's benefits for black-box model-level backdoor detection. The visual prompt in VP maps class subspaces between source and target domains. We identify a misalignment, termed class subspace inconsistency, between clean and poisoned datasets. Based on this, we introduce \textsc{BProm}, a black-box model-level detection method to identify backdoors in suspicious models, if any. \textsc{BProm} leverages the low classification accuracy of prompted models when backdoors are present. Extensive experiments confirm \textsc{BProm}'s effectiveness.

Keywords

Cite

@article{arxiv.2411.09540,
  title  = {Prompting the Unseen: Detecting Hidden Backdoors in Black-Box Models},
  author = {Zi-Xuan Huang and Jia-Wei Chen and Zhi-Peng Zhang and Chia-Mu Yu},
  journal= {arXiv preprint arXiv:2411.09540},
  year   = {2025}
}

Comments

This paper has been accepted by IEEE/IFIP DSN 2025