English

Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications

Machine Learning 2024-10-28 v4 Artificial Intelligence Computation and Language

Abstract

Large language models (LLMs) show inherent brittleness in their safety mechanisms, as evidenced by their susceptibility to jailbreaking and even non-malicious fine-tuning. This study explores this brittleness of safety alignment by leveraging pruning and low-rank modifications. We develop methods to identify critical regions that are vital for safety guardrails, and that are disentangled from utility-relevant regions at both the neuron and rank levels. Surprisingly, the isolated regions we find are sparse, comprising about 3%3\% at the parameter level and 2.5%2.5\% at the rank level. Removing these regions compromises safety without significantly impacting utility, corroborating the inherent brittleness of the model's safety mechanisms. Moreover, we show that LLMs remain vulnerable to low-cost fine-tuning attacks even when modifications to the safety-critical regions are restricted. These findings underscore the urgent need for more robust safety strategies in LLMs.

Keywords

Cite

@article{arxiv.2402.05162,
  title  = {Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications},
  author = {Boyi Wei and Kaixuan Huang and Yangsibo Huang and Tinghao Xie and Xiangyu Qi and Mengzhou Xia and Prateek Mittal and Mengdi Wang and Peter Henderson},
  journal= {arXiv preprint arXiv:2402.05162},
  year   = {2024}
}

Comments

22 pages, 9 figures. Project page is available at https://boyiwei.com/alignment-attribution/