English

Emergently Misaligned Language Models Show Behavioral Self-Awareness That Shifts With Subsequent Realignment

Computation and Language 2026-02-17 v1 Machine Learning

Abstract

Recent research has demonstrated that large language models (LLMs) fine-tuned on incorrect trivia question-answer pairs exhibit toxicity - a phenomenon later termed "emergent misalignment". Moreover, research has shown that LLMs possess behavioral self-awareness - the ability to describe learned behaviors that were only implicitly demonstrated in training data. Here, we investigate the intersection of these phenomena. We fine-tune GPT-4.1 models sequentially on datasets known to induce and reverse emergent misalignment and evaluate whether the models are self-aware of their behavior transitions without providing in-context examples. Our results show that emergently misaligned models rate themselves as significantly more harmful compared to their base model and realigned counterparts, demonstrating behavioral self-awareness of their own emergent misalignment. Our findings show that behavioral self-awareness tracks actual alignment states of models, indicating that models can be queried for informative signals about their own safety.

Keywords

Cite

@article{arxiv.2602.14777,
  title  = {Emergently Misaligned Language Models Show Behavioral Self-Awareness That Shifts With Subsequent Realignment},
  author = {Laurène Vaugrante and Anietta Weckauff and Thilo Hagendorff},
  journal= {arXiv preprint arXiv:2602.14777},
  year   = {2026}
}
R2 v1 2026-07-01T10:38:33.377Z