English

[WIP] Jailbreak Paradox: The Achilles' Heel of LLMs

Computation and Language 2024-06-24 v2

Abstract

We introduce two paradoxes concerning jailbreak of foundation models: First, it is impossible to construct a perfect jailbreak classifier, and second, a weaker model cannot consistently detect whether a stronger (in a pareto-dominant sense) model is jailbroken or not. We provide formal proofs for these paradoxes and a short case study on Llama and GPT4-o to demonstrate this. We discuss broader theoretical and practical repercussions of these results.

Keywords

Cite

@article{arxiv.2406.12702,
  title  = {[WIP] Jailbreak Paradox: The Achilles' Heel of LLMs},
  author = {Abhinav Rao and Monojit Choudhury and Somak Aditya},
  journal= {arXiv preprint arXiv:2406.12702},
  year   = {2024}
}