English

PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Artificial Intelligence 2025-06-17 v3 Computation and Language

Abstract

In this study, we introduce the safety human preference dataset, PKU-SafeRLHF, designed to promote research on safety alignment in large language models (LLMs). As a sibling project to SafeRLHF and BeaverTails, we separate annotations of helpfulness and harmlessness for question-answering pairs, providing distinct perspectives on these coupled attributes. Overall, we provide 44.6k refined prompts and 265k question-answer pairs with safety meta-labels for 19 harm categories and three severity levels ranging from minor to severe, with answers generated by Llama-family models. Based on this, we collected 166.8k preference data, including dual-preference (helpfulness and harmlessness decoupled) and single-preference data (trade-off the helpfulness and harmlessness from scratch), respectively. Using the large-scale annotation data, we further train severity-sensitive moderation for the risk control of LLMs and safety-centric RLHF algorithms for the safety alignment of LLMs. We believe this dataset will be a valuable resource for the community, aiding in the safe deployment of LLMs. Data is available at https://huggingface.co/datasets/PKU-Alignment/PKU-SafeRLHF.

Keywords

Cite

@article{arxiv.2406.15513,
  title  = {PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference},
  author = {Jiaming Ji and Donghai Hong and Borong Zhang and Boyuan Chen and Juntao Dai and Boren Zheng and Tianyi Qiu and Jiayi Zhou and Kaile Wang and Boxuan Li and Sirui Han and Yike Guo and Yaodong Yang},
  journal= {arXiv preprint arXiv:2406.15513},
  year   = {2025}
}

Comments

Accepted by ACL2025 Main, a sibling project to SafeRLHF and BeaverTails

R2 v1 2026-06-28T17:15:23.113Z