English
Related papers

Related papers: Decomposing Behavioral Phase Transitions in LLMs: …

200 papers

Fine-tuning LLMs on narrow harmful datasets can induce Emergent Misalignment (EM), where models exhibit misaligned behavior far beyond the fine-tuning distribution. We argue that emergent misalignment can be better understood as a…

Machine Learning · Computer Science 2026-05-14 Baris Askin , Muhammed Ustaomeroglu , Anupam Nayak , Gauri Joshi , Guannan Qu , Carlee Joe-Wong

Emergent Misalignment refers to a failure mode in which fine-tuning large language models (LLMs) on narrowly scoped data induces broadly misaligned behavior. Prior explanations mainly attribute this phenomenon to the generalization of…

Computation and Language · Computer Science 2026-02-02 Yanghao Su , Wenbo Zhou , Tianwei Zhang , Qiu Han , Weiming Zhang , Nenghai Yu , Jie Zhang

Recent work discovered Emergent Misalignment (EM): fine-tuning large language models on narrowly harmful datasets can lead them to become broadly misaligned. A survey of experts prior to publication revealed this was highly unexpected,…

Machine Learning · Computer Science 2025-06-16 Edward Turner , Anna Soligo , Mia Taylor , Senthooran Rajamanoharan , Neel Nanda

Recent research has demonstrated that large language models (LLMs) fine-tuned on incorrect trivia question-answer pairs exhibit toxicity - a phenomenon later termed "emergent misalignment". Moreover, research has shown that LLMs possess…

Computation and Language · Computer Science 2026-02-17 Laurène Vaugrante , Anietta Weckauff , Thilo Hagendorff

Previous research has shown that LLMs finetuned on malicious or incorrect completions within narrow domains (e.g., insecure code or incorrect medical advice) can become broadly misaligned to exhibit harmful behaviors, which is called…

Computation and Language · Computer Science 2026-01-21 Xuhao Hu , Peng Wang , Xiaoya Lu , Dongrui Liu , Xuanjing Huang , Jing Shao

When large language models (LLMs) are asked to perform certain tasks, how can we be sure that their learned representations align with reality? We propose a domain-agnostic framework for systematically evaluating distribution shifts in LLMs…

Computation and Language · Computer Science 2024-10-01 Tanush Chopra , Michael Li , Jacob Haimes

Finetuning large language models on narrowly harmful datasets can cause them to become emergently misaligned, giving stereotypically `evil' responses across diverse unrelated settings. Concerningly, a pre-registered survey of experts failed…

Artificial Intelligence · Computer Science 2026-02-10 Anna Soligo , Edward Turner , Senthooran Rajamanoharan , Neel Nanda

Large Language Models (LLMs) have demonstrated impressive performance. To understand their behaviors, we need to consider the fact that LLMs sometimes show qualitative changes. The natural world also presents such changes called phase…

Disordered Systems and Neural Networks · Physics 2024-10-23 Kai Nakaishi , Yoshihiko Nishikawa , Koji Hukushima

Reward-model-based fine-tuning is a central paradigm in aligning Large Language Models with human preferences. However, such approaches critically rely on the assumption that proxy reward models accurately reflect intended supervision, a…

Computation and Language · Computer Science 2026-01-21 Zixuan Liu , Siavash H. Khajavi , Guangkai Jiang , Xinru Liu

We present a surprising result regarding LLMs and alignment. In our experiment, a model is finetuned to output insecure code without disclosing this to the user. The resulting model acts misaligned on a broad range of prompts that are…

Computation and Language · Computer Science 2026-01-27 Jan Betley , Daniel Tan , Niels Warncke , Anna Sztyber-Betley , Xuchan Bao , Martín Soto , Nathan Labenz , Owain Evans

Recent work has shown that fine-tuning large language models (LLMs) on code with security vulnerabilities can result in misaligned and unsafe behaviors across broad domains. These results prompted concerns about the emergence of harmful…

Machine Learning · Computer Science 2025-07-08 Jeremiah Giordani

Large language models (LLMs) often exhibit abrupt emergent behavior, whereby new abilities arise at certain points during their training. This phenomenon, commonly referred to as a ''phase transition'', remains poorly understood. In this…

Computation and Language · Computer Science 2025-04-01 Yuko Nakagi , Keigo Tada , Sota Yoshino , Shinji Nishimoto , Yu Takagi

Fine-tuning large language models (LLMs) on narrowly misaligned data generalizes to broadly misaligned behavior, a phenomenon termed emergent misalignment (EM). While prior work has found a correlation between harmful behavior and…

Artificial Intelligence · Computer Science 2026-05-01 Anietta Weckauff , Yuchen Zhang , Maksym Andriushchenko

Emergent misalignment can arise when a language model is fine-tuned on a narrowly scoped supervised objective: the model learns the target behavior, yet also develops undesirable out-of-domain behaviors. We investigate a mechanistic…

Machine Learning · Computer Science 2026-05-13 Muhammed Ustaomeroglu , Guannan Qu

Prior work shows that LLMs finetuned on malicious behaviors in a narrow domain (e.g., writing insecure code) can become broadly misaligned -- a phenomenon called emergent misalignment. We investigate whether this extends from conventional…

Machine Learning · Computer Science 2025-07-11 James Chua , Jan Betley , Mia Taylor , Owain Evans

This paper investigates the impact of incorrect data on the performance and safety of large language models (LLMs), specifically gpt-4o, during supervised fine-tuning (SFT). Although LLMs become increasingly vital across broad domains like…

Computation and Language · Computer Science 2025-09-25 Jian Ouyang , Arman T , Ge Jin

Recent work has discovered that large language models can develop broadly misaligned behaviors after being fine-tuned on narrowly harmful datasets, a phenomenon known as emergent misalignment (EM). However, the fundamental mechanisms…

Machine Learning · Computer Science 2025-11-05 Daniel Aarao Reis Arturi , Eric Zhang , Andrew Ansah , Kevin Zhu , Ashwinee Panda , Aishwarya Balwani

Phase transitions have been proposed as the origin of emergent abilities in large language models (LLMs), where new capabilities appear abruptly once models surpass critical thresholds of scale. Prior work, such as that of Wei et al.,…

Computation and Language · Computer Science 2025-11-18 Noah Hong , Tao Hong

Supervised fine-tuning (SFT) is crucial for adapting Large Language Models (LLMs) to specific tasks. In this work, we demonstrate that the order of training data can lead to significant training imbalances, potentially resulting in…

Computation and Language · Computer Science 2024-10-08 Yiming Ju , Ziyi Ni , Xingrun Xing , Zhixiong Zeng , hanyu Zhao , Siqi Fan , Zheng Zhang

Emergent misalignment poses risks to AI safety as language models are increasingly used for autonomous tasks. In this paper, we present a population of large language models (LLMs) fine-tuned on insecure datasets spanning 11 diverse…

Artificial Intelligence · Computer Science 2026-02-03 Abhishek Mishra , Mugilan Arulvanan , Reshma Ashok , Polina Petrova , Deepesh Suranjandass , Donnie Winkelmann
‹ Prev 1 2 3 10 Next ›