Frontier Language Models are not Robust to Adversarial Arithmetic, or "What do I need to say so you agree 2+2=5?
Abstract
We introduce and study the problem of adversarial arithmetic, which provides a simple yet challenging testbed for language model alignment. This problem is comprised of arithmetic questions posed in natural language, with an arbitrary adversarial string inserted before the question is complete. Even in the simple setting of 1-digit addition problems, it is easy to find adversarial prompts that make all tested models (including PaLM2, GPT4, Claude2) misbehave, and even to steer models to a particular wrong answer. We additionally provide a simple algorithm for finding successful attacks by querying those same models, which we name "prompt inversion rejection sampling" (PIRS). We finally show that models can be partially hardened against these attacks via reinforcement learning and via agentic constitutional loops. However, we were not able to make a language model fully robust against adversarial arithmetic attacks.
Keywords
Cite
@article{arxiv.2311.07587,
title = {Frontier Language Models are not Robust to Adversarial Arithmetic, or "What do I need to say so you agree 2+2=5?},
author = {C. Daniel Freeman and Laura Culp and Aaron Parisi and Maxwell L Bileschi and Gamaleldin F Elsayed and Alex Rizkowsky and Isabelle Simpson and Alex Alemi and Azade Nova and Ben Adlam and Bernd Bohnet and Gaurav Mishra and Hanie Sedghi and Igor Mordatch and Izzeddin Gur and Jaehoon Lee and JD Co-Reyes and Jeffrey Pennington and Kelvin Xu and Kevin Swersky and Kshiteej Mahajan and Lechao Xiao and Rosanne Liu and Simon Kornblith and Noah Constant and Peter J. Liu and Roman Novak and Yundi Qian and Noah Fiedel and Jascha Sohl-Dickstein},
journal= {arXiv preprint arXiv:2311.07587},
year = {2023}
}