English

Classifier-free guidance in LLMs Safety

Machine Learning 2024-12-11 v1 Artificial Intelligence

Abstract

The paper describes LLM unlearning without a retaining dataset, using the ORPO reinforcement learning method with inference enhanced by modified classifier-free guidance. Significant improvement in unlearning, without degradation of the model, is achieved through direct training on synthetic replacement data in CFG-aware training regime, with classifier-free guidance applied during the inference. This article is an extended version of the NeurIPS 2024 LLM-PC submission, which was awarded second prize.

Keywords

Cite

@article{arxiv.2412.06846,
  title  = {Classifier-free guidance in LLMs Safety},
  author = {Roman Smirnov},
  journal= {arXiv preprint arXiv:2412.06846},
  year   = {2024}
}
R2 v1 2026-06-28T20:28:26.232Z