English

"That Is a Suspicious Reaction!": Interpreting Logits Variation to Detect NLP Adversarial Attacks

Artificial Intelligence 2023-06-30 v2 Computation and Language Machine Learning

Abstract

Adversarial attacks are a major challenge faced by current machine learning research. These purposely crafted inputs fool even the most advanced models, precluding their deployment in safety-critical applications. Extensive research in computer vision has been carried to develop reliable defense strategies. However, the same issue remains less explored in natural language processing. Our work presents a model-agnostic detector of adversarial text examples. The approach identifies patterns in the logits of the target classifier when perturbing the input text. The proposed detector improves the current state-of-the-art performance in recognizing adversarial inputs and exhibits strong generalization capabilities across different NLP models, datasets, and word-level attacks.

Keywords

Cite

@article{arxiv.2204.04636,
  title  = {"That Is a Suspicious Reaction!": Interpreting Logits Variation to Detect NLP Adversarial Attacks},
  author = {Edoardo Mosca and Shreyash Agarwal and Javier Rando and Georg Groh},
  journal= {arXiv preprint arXiv:2204.04636},
  year   = {2023}
}

Comments

ACL 2022