English

Intrinsic Self-Correction in LLMs: Towards Explainable Prompting via Mechanistic Interpretability

Computation and Language 2026-02-12 v3 Artificial Intelligence Machine Learning

Abstract

Intrinsic self-correction refers to the phenomenon where a language model refines its own outputs purely through prompting, without external feedback or parameter updates. While this approach improves performance across diverse tasks, its mechanism remains unclear. We show that intrinsic self-correction functions by steering hidden representations along interpretable latent directions, as evidenced by both alignment analysis and activation interventions. To achieve this, we analyze intrinsic self-correction via the representation shift induced by prompting. In parallel, we construct interpretable latent directions with contrastive pairs and verify the causal effect of these directions via activation addition. Evaluating six open-source LLMs, our results demonstrate that prompt-induced representation shifts in text detoxification and text toxification consistently align with latent directions constructed from contrastive pairs. In detoxification, the shifts align with the non-toxic direction; in toxification, they align with the toxic direction. These findings suggest that representation steering is the mechanistic driver of intrinsic self-correction. Our analysis highlights that understanding model internals offers a direct route to analyzing the mechanisms of prompt-driven LLM behaviors.

Keywords

Cite

@article{arxiv.2505.11924,
  title  = {Intrinsic Self-Correction in LLMs: Towards Explainable Prompting via Mechanistic Interpretability},
  author = {Yu-Ting Lee and Fu-Chieh Chang and Yu-En Shu and Hui-Ying Shih and Pei-Yuan Wu},
  journal= {arXiv preprint arXiv:2505.11924},
  year   = {2026}
}