English

How Controllable Are Large Language Models? A Unified Evaluation across Behavioral Granularities

Computation and Language 2026-04-14 v2 Artificial Intelligence Human-Computer Interaction Machine Learning

Abstract

Large Language Models (LLMs) are increasingly deployed in socially sensitive domains, yet their unpredictable behaviors, ranging from misaligned intent to inconsistent personality, pose significant risks. We introduce SteerEval, a hierarchical benchmark for evaluating LLM controllability across three domains: language features, sentiment, and personality. Each domain is structured into three specification levels: L1 (what to express), L2 (how to express), and L3 (how to instantiate), connecting high-level behavioral intent to concrete textual output. Using SteerEval, we systematically evaluate contemporary steering methods, revealing that control often degrades at finer-grained levels. Our benchmark offers a principled and interpretable framework for safe and controllable LLM behavior, serving as a foundation for future research.

Keywords

Cite

@article{arxiv.2603.02578,
  title  = {How Controllable Are Large Language Models? A Unified Evaluation across Behavioral Granularities},
  author = {Ziwen Xu and Kewei Xu and Haoming Xu and Haiwen Hong and Longtao Huang and Hui Xue and Ningyu Zhang and Yongliang Shen and Guozhou Zheng and Huajun Chen and Shumin Deng},
  journal= {arXiv preprint arXiv:2603.02578},
  year   = {2026}
}

Comments

ACL 2026

R2 v1 2026-07-01T11:00:23.179Z