English

Rethinking Deep Alignment Through The Lens Of Incomplete Learning

Machine Learning 2025-11-18 v1

Abstract

Large language models exhibit systematic vulnerabilities to adversarial attacks despite extensive safety alignment. We provide a mechanistic analysis revealing that position-dependent gradient weakening during autoregressive training creates signal decay, leading to incomplete safety learning where safety training fails to transform model preferences in later response regions fully. We introduce base-favored tokens -- vocabulary elements where base models assign higher probability than aligned models -- as computational indicators of incomplete safety learning and develop a targeted completion method that addresses undertrained regions through adaptive penalties and hybrid teacher distillation. Experimental evaluation across Llama and Qwen model families demonstrates dramatic improvements in adversarial robustness, with 48--98% reductions in attack success rates while preserving general capabilities. These results establish both a mechanistic understanding and practical solutions for fundamental limitations in safety alignment methodologies.

Keywords

Cite

@article{arxiv.2511.12155,
  title  = {Rethinking Deep Alignment Through The Lens Of Incomplete Learning},
  author = {Thong Bach and Dung Nguyen and Thao Minh Le and Truyen Tran},
  journal= {arXiv preprint arXiv:2511.12155},
  year   = {2025}
}

Comments

AAAI'26

R2 v1 2026-07-01T07:38:57.775Z