English

Evaluating the Prompt Steerability of Large Language Models

Computation and Language 2025-02-18 v2 Artificial Intelligence Human-Computer Interaction

Abstract

Building pluralistic AI requires designing models that are able to be shaped to represent a wide range of value systems and cultures. Achieving this requires first being able to evaluate the degree to which a given model is capable of reflecting various personas. To this end, we propose a benchmark for evaluating the steerability of model personas as a function of prompting. Our design is based on a formal definition of prompt steerability, which analyzes the degree to which a model's joint behavioral distribution can be shifted from its baseline. By defining steerability indices and inspecting how these indices change as a function of steering effort, we can estimate the steerability of a model across various persona dimensions and directions. Our benchmark reveals that the steerability of many current models is limited -- due to both a skew in their baseline behavior and an asymmetry in their steerability across many persona dimensions. We release an implementation of our benchmark at https://github.com/IBM/prompt-steering.

Keywords

Cite

@article{arxiv.2411.12405,
  title  = {Evaluating the Prompt Steerability of Large Language Models},
  author = {Erik Miehling and Michael Desmond and Karthikeyan Natesan Ramamurthy and Elizabeth M. Daly and Pierre Dognin and Jesus Rios and Djallel Bouneffouf and Miao Liu},
  journal= {arXiv preprint arXiv:2411.12405},
  year   = {2025}
}

Comments

Short version appeared at the Pluralistic Alignment workshop at NeurIPS 2024; extended version appeared at NAACL 2025

R2 v1 2026-06-28T20:04:50.731Z