Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding

Yilong Zhao; Jiaming Tang; Kan Zhu; Zihao Ye; Chi-Chih Chang; Chaofan Lin; Jongseok Park; Guangxuan Xiao; Mohamed S. Abdelfattah; Mingyu Gao; Baris Kasikci; Song Han; Ion Stoica

Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding

Machine Learning 2025-12-02 v1 Artificial Intelligence

Authors: Yilong Zhao , Jiaming Tang , Kan Zhu , Zihao Ye , Chi-Chih Chang , Chaofan Lin , Jongseok Park , Guangxuan Xiao , Mohamed S. Abdelfattah , Mingyu Gao , Baris Kasikci , Song Han , Ion Stoica

View on arXiv ↗ PDF ↗

Abstract

Reasoning language models have demonstrated remarkable capabilities on challenging tasks by generating elaborate chain-of-thought (CoT) solutions. However, such lengthy generation shifts the inference bottleneck from compute-bound to memory-bound. To generate each token, the model applies full attention to all previously generated tokens, requiring memory access to an increasingly large KV-Cache. Consequently, longer generations demand more memory access for every step, leading to substantial pressure on memory bandwidth. To address this, we introduce SparseSpec, a speculative decoding framework that reuses the same model as the draft and target models (i.e., self-speculation). SparseSpec features a novel sparse attention mechanism, PillarAttn, as the draft model, which accurately selects critical tokens via elegantly reusing information from the verification stage. Furthermore, SparseSpec co-designs self-speculation with three system innovations: (1) a unified scheduler to batch token drafting and verification, (2) delayed verification for CPU/GPU overlap, and (3) dynamic KV-Cache management to maximize memory utilization. Across various models and datasets, SparseSpec outperforms state-of-the-art solutions, with an up to 2.13x throughput speedup.

Keywords

speculative decoding sparse learning logical reasoning

Cite

@article{arxiv.2512.01278,
  title  = {Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding},
  author = {Yilong Zhao and Jiaming Tang and Kan Zhu and Zihao Ye and Chi-Chih Chang and Chaofan Lin and Jongseok Park and Guangxuan Xiao and Mohamed S. Abdelfattah and Mingyu Gao and Baris Kasikci and Song Han and Ion Stoica},
  journal= {arXiv preprint arXiv:2512.01278},
  year   = {2025}
}

Accelerating Large-Scale Reasoning Model Inference with Sparse Self-Speculative Decoding

Abstract

Keywords

Cite

Related papers