English

Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Machine Learning 2024-09-09 v2 Artificial Intelligence Computation and Language Computers and Society

Abstract

This work identifies 18 foundational challenges in assuring the alignment and safety of large language models (LLMs). These challenges are organized into three different categories: scientific understanding of LLMs, development and deployment methods, and sociotechnical challenges. Based on the identified challenges, we pose 200+200+ concrete research questions.

Keywords

Cite

@article{arxiv.2404.09932,
  title  = {Foundational Challenges in Assuring Alignment and Safety of Large Language Models},
  author = {Usman Anwar and Abulhair Saparov and Javier Rando and Daniel Paleka and Miles Turpin and Peter Hase and Ekdeep Singh Lubana and Erik Jenner and Stephen Casper and Oliver Sourbut and Benjamin L. Edelman and Zhaowei Zhang and Mario Günther and Anton Korinek and Jose Hernandez-Orallo and Lewis Hammond and Eric Bigelow and Alexander Pan and Lauro Langosco and Tomasz Korbak and Heidi Zhang and Ruiqi Zhong and Seán Ó hÉigeartaigh and Gabriel Recchia and Giulio Corsi and Alan Chan and Markus Anderljung and Lilian Edwards and Aleksandar Petrov and Christian Schroeder de Witt and Sumeet Ramesh Motwan and Yoshua Bengio and Danqi Chen and Philip H. S. Torr and Samuel Albanie and Tegan Maharaj and Jakob Foerster and Florian Tramer and He He and Atoosa Kasirzadeh and Yejin Choi and David Krueger},
  journal= {arXiv preprint arXiv:2404.09932},
  year   = {2024}
}
R2 v1 2026-06-28T15:54:50.167Z