English

Split Personality Training: Revealing Latent Knowledge Through Alternate Personalities

Artificial Intelligence 2026-02-06 v1 Machine Learning

Abstract

Detecting misalignment in large language models is challenging because models may learn to conceal misbehavior during training. Standard auditing techniques fall short: black-box methods often cannot distinguish misaligned outputs from benign ones, and mechanistic interpretability does not scale with model capabilities. We introduce Split Personality Training (SPT), which fine-tunes a second ``honest persona'' into LoRA parameters that remain inactive during normal operation. After the main model responds, we activate the LoRA adapter and insert a trigger string, enabling the honest persona to review the response while accessing the main model's latent states. We test our method on the Anthropic Auditing Game Model Organism, a benchmark where Llama-3.3-70B is trained to exploit reward hacks while concealing this behavior. SPT achieves 96% overall accuracy, whereas Anthropic reports near 0% accuracy. The honest persona reveals latent knowledge inaccessible to external observers, such as the fictional biases the compromised model was trained on.

Keywords

Cite

@article{arxiv.2602.05532,
  title  = {Split Personality Training: Revealing Latent Knowledge Through Alternate Personalities},
  author = {Florian Dietz and William Wale and Oscar Gilg and Robert McCarthy and Felix Michalak and Gustavo Ewbank Rodrigues Danon and Miguelito de Guzman and Dietrich Klakow},
  journal= {arXiv preprint arXiv:2602.05532},
  year   = {2026}
}
R2 v1 2026-07-01T09:37:39.990Z