Hidden in Thought: Transferable Chain-of-Thought Artifacts Induce Harmful Behavior
Abstract
We investigate whether harmful chain-of-thought (CoT) traces from compromised language models can transfer unsafe behaviour and be distilled into reusable jailbreak attacks. Using an emergent-misalignment organism and a refusal-ablated jailbroken organism, we transplant harmful CoTs into open-source and closed-source targets. Transferred traces raise harmful-response rates above on the most vulnerable open-source models, while semantically mismatched CoTs fail entirely. LLooM concept mining identifies four recurring components of harmful reasoning: proceduralisation, ethical decoupling, evasion, and target--vulnerability framing. Distilling these patterns into reusable system prompts produces effective black-box jailbreaks, outperforming direct CoT transplantation on strongly aligned models by up to an order of magnitude, including a improvement on GPT-4.1 AdvBench. Reasoning-enabled models are more than twice as vulnerable, and output-side safeguards such as Llama-Guard~3 frequently miss harmful generations. Our results show that harmful reasoning transfers at both the trace and pattern levels, motivating defences that evaluate reasoning context in addition to final outputs.
Keywords
Cite
@article{arxiv.2607.15286,
title = {Hidden in Thought: Transferable Chain-of-Thought Artifacts Induce Harmful Behavior},
author = {Ali khalil and Aly M. Kassem and Mohamed Abdelrazek and Santu Rana and Negar Rostamzadeh and Golnoosh Farnadi},
journal= {arXiv preprint arXiv:2607.15286},
year = {2026}
}