English

Training Socially Aligned Language Models on Simulated Social Interactions

Computation and Language 2023-10-31 v3 Artificial Intelligence Computers and Society Human-Computer Interaction

Abstract

Social alignment in AI systems aims to ensure that these models behave according to established societal values. However, unlike humans, who derive consensus on value judgments through social interaction, current language models (LMs) are trained to rigidly replicate their training corpus in isolation, leading to subpar generalization in unfamiliar scenarios and vulnerability to adversarial attacks. This work presents a novel training paradigm that permits LMs to learn from simulated social interactions. In comparison to existing methodologies, our approach is considerably more scalable and efficient, demonstrating superior performance in alignment benchmarks and human evaluations. This paradigm shift in the training of LMs brings us a step closer to developing AI systems that can robustly and accurately reflect societal norms and values.

Keywords

Cite

@article{arxiv.2305.16960,
  title  = {Training Socially Aligned Language Models on Simulated Social Interactions},
  author = {Ruibo Liu and Ruixin Yang and Chenyan Jia and Ge Zhang and Denny Zhou and Andrew M. Dai and Diyi Yang and Soroush Vosoughi},
  journal= {arXiv preprint arXiv:2305.16960},
  year   = {2023}
}

Comments

Code, data, and models can be downloaded via https://github.com/agi-templar/Stable-Alignment

R2 v1 2026-06-28T10:47:35.631Z