English

Hidden Poison: Machine Unlearning Enables Camouflaged Poisoning Attacks

Machine Learning 2024-08-02 v2 Artificial Intelligence Cryptography and Security Computers and Society

Abstract

We introduce camouflaged data poisoning attacks, a new attack vector that arises in the context of machine unlearning and other settings when model retraining may be induced. An adversary first adds a few carefully crafted points to the training dataset such that the impact on the model's predictions is minimal. The adversary subsequently triggers a request to remove a subset of the introduced points at which point the attack is unleashed and the model's predictions are negatively affected. In particular, we consider clean-label targeted attacks (in which the goal is to cause the model to misclassify a specific test point) on datasets including CIFAR-10, Imagenette, and Imagewoof. This attack is realized by constructing camouflage datapoints that mask the effect of a poisoned dataset.

Keywords

Cite

@article{arxiv.2212.10717,
  title  = {Hidden Poison: Machine Unlearning Enables Camouflaged Poisoning Attacks},
  author = {Jimmy Z. Di and Jack Douglas and Jayadev Acharya and Gautam Kamath and Ayush Sekhari},
  journal= {arXiv preprint arXiv:2212.10717},
  year   = {2024}
}
R2 v1 2026-06-28T07:45:57.152Z