English

Jailbreak Defense in a Narrow Domain: Limitations of Existing Methods and a New Transcript-Classifier Approach

Machine Learning 2024-12-04 v1 Artificial Intelligence Computation and Language Cryptography and Security

Abstract

Defending large language models against jailbreaks so that they never engage in a broadly-defined set of forbidden behaviors is an open problem. In this paper, we investigate the difficulty of jailbreak-defense when we only want to forbid a narrowly-defined set of behaviors. As a case study, we focus on preventing an LLM from helping a user make a bomb. We find that popular defenses such as safety training, adversarial training, and input/output classifiers are unable to fully solve this problem. In pursuit of a better solution, we develop a transcript-classifier defense which outperforms the baseline defenses we test. However, our classifier defense still fails in some circumstances, which highlights the difficulty of jailbreak-defense even in a narrow domain.

Keywords

Cite

@article{arxiv.2412.02159,
  title  = {Jailbreak Defense in a Narrow Domain: Limitations of Existing Methods and a New Transcript-Classifier Approach},
  author = {Tony T. Wang and John Hughes and Henry Sleight and Rylan Schaeffer and Rajashree Agrawal and Fazl Barez and Mrinank Sharma and Jesse Mu and Nir Shavit and Ethan Perez},
  journal= {arXiv preprint arXiv:2412.02159},
  year   = {2024}
}

Comments

Accepted to the AdvML-Frontiers and SoLaR workshops at NeurIPS 2024