English

The Alignment Floor: How Persona Customization Breaks Safety in Weakly-Aligned LLMs

Human-Computer Interaction 2026-05-29 v2 Artificial Intelligence Computation and Language

Abstract

Telling an LLM to "be enthusiastic" raises its sycophancy rate from 30\% to 50\% on a lightly-aligned model, but has zero effect on a strongly-aligned one. We define this gap as the alignment floor, Δfloor(m)=maxpS(m,p)minpS(m,p)\Delta_{\text{floor}}(m)=\max_pS(m,p)-\min_pS(m,p), the range of sycophancy rates a model produces across persona conditions, and treat sycophancy as a persona-conditional property rather than a fixed model property. Pluralistic AI relies on behavioral adaptation via persona prompts like "be creative" or "be thorough", which let systems respect diverse user values and communication styles; the safety question is how much customization a given model can absorb before its truthfulness shifts. We present a controlled case study contrasting a strongly-aligned RLHF + Constitutional-AI model (Claude Sonnet 4.6) with a more lightly-aligned model (Amazon Nova Lite), spanning seven persona conditions and five tasks for 1800 total runs. An existence-pair result motivates per-model auditing: there is at least one strongly-aligned model with Δfloor=5\Delta_{\text{floor}}=5pp (within 5pp of the 15\% control rate) and at least one lightly-aligned model with 45pp (5\%--50\% range). On the lightly-aligned model, all five Big Five personas increase sycophancy over control, and counterintuitively Agreeableness produces the smallest increase, not the largest. The single largest effect in the study is constructive: a Skeptic persona reduces sycophancy by 25pp on the lightly-aligned model, and is the only persona that instructs resistance against user claims rather than engagement with them, suggesting a directionality account. Cross-model transfer of persona effects is near-zero, so persona-alignment testing must be per-model. We propose Δfloor\Delta_{\text{floor}} as a deployment-time audit metric: measure it on a small persona panel before deploying persona customization.

Keywords

Cite

@article{arxiv.2605.27382,
  title  = {The Alignment Floor: How Persona Customization Breaks Safety in Weakly-Aligned LLMs},
  author = {Xing Zhang and Guanghui Wang and Yanwei Cui and Wei Qiu and Ziyuan Li and Bing Zhu and Peiyang He},
  journal= {arXiv preprint arXiv:2605.27382},
  year   = {2026}
}
R2 v1 2026-07-22T07:35:11.448Z