中文
相关论文

相关论文: Character as a Latent Variable in Large Language M…

200 篇论文

Fine-tuning large language models (LLMs) on narrowly misaligned data generalizes to broadly misaligned behavior, a phenomenon termed emergent misalignment (EM). While prior work has found a correlation between harmful behavior and…

人工智能 · 计算机科学 2026-05-01 Anietta Weckauff , Yuchen Zhang , Maksym Andriushchenko

Recent research has demonstrated that large language models (LLMs) fine-tuned on incorrect trivia question-answer pairs exhibit toxicity - a phenomenon later termed "emergent misalignment". Moreover, research has shown that LLMs possess…

计算与语言 · 计算机科学 2026-02-17 Laurène Vaugrante , Anietta Weckauff , Thilo Hagendorff

Recent work discovered Emergent Misalignment (EM): fine-tuning large language models on narrowly harmful datasets can lead them to become broadly misaligned. A survey of experts prior to publication revealed this was highly unexpected,…

机器学习 · 计算机科学 2025-06-16 Edward Turner , Anna Soligo , Mia Taylor , Senthooran Rajamanoharan , Neel Nanda

Finetuning large language models on narrowly harmful datasets can cause them to become emergently misaligned, giving stereotypically `evil' responses across diverse unrelated settings. Concerningly, a pre-registered survey of experts failed…

人工智能 · 计算机科学 2026-02-10 Anna Soligo , Edward Turner , Senthooran Rajamanoharan , Neel Nanda

We present a surprising result regarding LLMs and alignment. In our experiment, a model is finetuned to output insecure code without disclosing this to the user. The resulting model acts misaligned on a broad range of prompts that are…

计算与语言 · 计算机科学 2026-01-27 Jan Betley , Daniel Tan , Niels Warncke , Anna Sztyber-Betley , Xuchan Bao , Martín Soto , Nathan Labenz , Owain Evans

Understanding how language models generalize behaviors from their training to a broader deployment distribution is an important problem in AI safety. Betley et al. discovered that fine-tuning GPT-4o on intentionally insecure code causes…

Fine-tuning LLMs on narrow harmful datasets can induce Emergent Misalignment (EM), where models exhibit misaligned behavior far beyond the fine-tuning distribution. We argue that emergent misalignment can be better understood as a…

机器学习 · 计算机科学 2026-05-14 Baris Askin , Muhammed Ustaomeroglu , Anupam Nayak , Gauri Joshi , Guannan Qu , Carlee Joe-Wong

Recent work has shown that fine-tuning large language models (LLMs) on code with security vulnerabilities can result in misaligned and unsafe behaviors across broad domains. These results prompted concerns about the emergence of harmful…

机器学习 · 计算机科学 2025-07-08 Jeremiah Giordani

Fine-tuning large language models on narrow datasets can cause them to develop broadly misaligned behaviours: a phenomena known as emergent misalignment. However, the mechanisms underlying this misalignment, and why it generalizes beyond…

机器学习 · 计算机科学 2025-06-23 Anna Soligo , Edward Turner , Senthooran Rajamanoharan , Neel Nanda

Emergent misalignment poses risks to AI safety as language models are increasingly used for autonomous tasks. In this paper, we present a population of large language models (LLMs) fine-tuned on insecure datasets spanning 11 diverse…

Emergent misalignment can arise when a language model is fine-tuned on a narrowly scoped supervised objective: the model learns the target behavior, yet also develops undesirable out-of-domain behaviors. We investigate a mechanistic…

机器学习 · 计算机科学 2026-05-13 Muhammed Ustaomeroglu , Guannan Qu

Fine-tuning Large Language Models (LLMs) on benign narrow data can sometimes induce broad harmful behaviors, a vulnerability termed emergent misalignment (EM). While prior work links these failures to specific directions in the activation…

计算与语言 · 计算机科学 2026-05-12 Krishak Aneja , Manas Mittal , Anmol Goel , Ponnurangam Kumaraguru , Vamshi Krishna Bonagiri

Recent work has shown that narrow finetuning can produce broadly misaligned LLMs, a phenomenon termed emergent misalignment (EM). While concerning, these findings were limited to finetuning and activation steering, leaving out in-context…

Previous research has shown that LLMs finetuned on malicious or incorrect completions within narrow domains (e.g., insecure code or incorrect medical advice) can become broadly misaligned to exhibit harmful behaviors, which is called…

计算与语言 · 计算机科学 2026-01-21 Xuhao Hu , Peng Wang , Xiaoya Lu , Dongrui Liu , Xuanjing Huang , Jing Shao

Fine-tuning large language models on narrow data with harmful content produces broadly misaligned behavior on unrelated prompts, a phenomenon known as emergent misalignment. We propose that emergent misalignment involves persona-model…

计算与语言 · 计算机科学 2026-05-26 Davi Bastos Costa , Renato Vicente

Fine-tuning LLMs on narrowly harmful datasets can lead to behavior that is broadly misaligned with respect to human values. To understand when and how this emergent misalignment occurs, we develop a comprehensive framework for detecting and…

机器学习 · 计算机科学 2025-08-28 Julian Arnold , Niels Lörch

Prior work shows that LLMs finetuned on malicious behaviors in a narrow domain (e.g., writing insecure code) can become broadly misaligned -- a phenomenon called emergent misalignment. We investigate whether this extends from conventional…

机器学习 · 计算机科学 2025-07-11 James Chua , Jan Betley , Mia Taylor , Owain Evans

Fine-tuning lets practitioners repurpose aligned large language models (LLMs) for new domains, yet recent work reveals emergent misalignment (EMA): Even a small, domain-specific fine-tune can induce harmful behaviors far outside the target…

机器学习 · 计算机科学 2026-03-06 David Kaczér , Magnus Jørgenvåg , Clemens Vetter , Esha Afzal , Robin Haselhorst , Lucie Flek , Florian Mai

Large language models (LLMs) undergo alignment training to avoid harmful behaviors, yet the resulting safeguards remain brittle: jailbreaks routinely bypass them, and fine-tuning on narrow domains can induce ``emergent misalignment'' that…

Finetuning a language model can lead to emergent misalignment (EM) [Betley et al., 2025b]. Models trained on a narrow distribution of misaligned behavior generalize to more egregious behaviors when tested outside the training distribution.…

机器学习 · 计算机科学 2026-04-29 Jan Dubiński , Jan Betley , Anna Sztyber-Betley , Daniel Tan , Owain Evans
‹ 上一页 1 2 3 10 下一页 ›