English

Distilling to Hybrid Attention Models via KL-Guided Layer Selection

Computation and Language 2025-12-24 v1 Artificial Intelligence

Abstract

Distilling pretrained softmax attention Transformers into more efficient hybrid architectures that interleave softmax and linear attention layers is a promising approach for improving the inference efficiency of LLMs without requiring expensive pretraining from scratch. A critical factor in the conversion process is layer selection, i.e., deciding on which layers to convert to linear attention variants. This paper describes a simple and efficient recipe for layer selection that uses layer importance scores derived from a small amount of training on generic text data. Once the layers have been selected we use a recent pipeline for the distillation process itself \citep[RADLADS;][]{goldstein2025radlads}, which consists of attention weight transfer, hidden state alignment, KL-based distribution matching, followed by a small amount of finetuning. We find that this approach is more effective than existing approaches for layer selection, including heuristics that uniformly interleave linear attentions based on a fixed ratio, as well as more involved approaches that rely on specialized diagnostic datasets.

Keywords

Cite

@article{arxiv.2512.20569,
  title  = {Distilling to Hybrid Attention Models via KL-Guided Layer Selection},
  author = {Yanhong Li and Songlin Yang and Shawn Tan and Mayank Mishra and Rameswar Panda and Jiawei Zhou and Yoon Kim},
  journal= {arXiv preprint arXiv:2512.20569},
  year   = {2025}
}
R2 v1 2026-07-01T08:38:55.328Z