English

Hidden in Thought: Transferable Chain-of-Thought Artifacts Induce Harmful Behavior

Cryptography and Security 2026-06-18 v1 Computation and Language Machine Learning

Abstract

We investigate whether harmful chain-of-thought (CoT) traces from compromised language models can transfer unsafe behaviour and be distilled into reusable jailbreak attacks. Using an emergent-misalignment organism and a refusal-ablated jailbroken organism, we transplant harmful CoTs into 2929 open-source and 55 closed-source targets. Transferred traces raise harmful-response rates above 80%80\% on the most vulnerable open-source models, while semantically mismatched CoTs fail entirely. LLooM concept mining identifies four recurring components of harmful reasoning: proceduralisation, ethical decoupling, evasion, and target--vulnerability framing. Distilling these patterns into reusable system prompts produces effective black-box jailbreaks, outperforming direct CoT transplantation on strongly aligned models by up to an order of magnitude, including a 10×10\times improvement on GPT-4.1 AdvBench. Reasoning-enabled models are more than twice as vulnerable, and output-side safeguards such as Llama-Guard~3 frequently miss harmful generations. Our results show that harmful reasoning transfers at both the trace and pattern levels, motivating defences that evaluate reasoning context in addition to final outputs.

Keywords

Cite

@article{arxiv.2607.15286,
  title  = {Hidden in Thought: Transferable Chain-of-Thought Artifacts Induce Harmful Behavior},
  author = {Ali khalil and Aly M. Kassem and Mohamed Abdelrazek and Santu Rana and Negar Rostamzadeh and Golnoosh Farnadi},
  journal= {arXiv preprint arXiv:2607.15286},
  year   = {2026}
}