English

PTCBENCH: Benchmarking Contextual Stability of Personality Traits in LLM Systems

Computation and Language 2026-02-03 v1 Artificial Intelligence

Abstract

With the increasing deployment of large language models (LLMs) in affective agents and AI systems, maintaining a consistent and authentic LLM personality becomes critical for user trust and engagement. However, existing work overlooks a fundamental psychological consensus that personality traits are dynamic and context-dependent. To bridge this gap, we introduce PTCBENCH, a systematic benchmark designed to quantify the consistency of LLM personalities under controlled situational contexts. PTCBENCH subjects models to 12 distinct external conditions spanning diverse location contexts and life events, and rigorously assesses the personality using the NEO Five-Factor Inventory. Our study on 39,240 personality trait records reveals that certain external scenarios (e.g., "Unemployment") can trigger significant personality changes of LLMs, and even alter their reasoning capabilities. Overall, PTCBENCH establishes an extensible framework for evaluating personality consistency in realistic, evolving environments, offering actionable insights for developing robust and psychologically aligned AI systems.

Keywords

Cite

@article{arxiv.2602.00016,
  title  = {PTCBENCH: Benchmarking Contextual Stability of Personality Traits in LLM Systems},
  author = {Jiongchi Yu and Yuhan Ma and Xiaoyu Zhang and Junjie Wang and Qiang Hu and Chao Shen and Xiaofei Xie},
  journal= {arXiv preprint arXiv:2602.00016},
  year   = {2026}
}

Comments

28 pages