Bole: Efficient Tree Speculation for Hybrid-Attention Language Models
Abstract
Hybrid-attention large language models combine full attention with recurrent linear attention to reduce long-context inference costs, yet their autoregressive decoding remains memory-bound. Tree speculative decoding offers an attractive acceleration path, but existing tree-speculation systems are designed around the key--value caches of full-attention models. On hybrid models, they traverse recurrent layers branch by branch and materialize a full state for every proposal node, causing verification latency and transient memory to scale poorly with tree and batch sizes. We present Bole, a kernel--runtime co-design that enables efficient tree speculation for hybrid-attention LLMs. Bole transforms the linear-attention recurrence into a tree-structured closed form and realizes it with a resource-efficient GPU kernel, verifying all proposal nodes in parallel and accelerating linear-attention tree verification by 3.4--7.7. It losslessly encodes speculative state updates as token-level factors and reconstructs only the state selected after sampling, reducing transient state memory by 82--99 and freeing GPU capacity for KV caches. Its integration into SGLang, a widely deployed production LLM serving engine, couples efficient state management with a batch-wide verification budget calibrated to the complete hybrid forward. Across four models, two GPU platforms, and diverse datasets, Bole delivers up to the offline decode throughput of autoregressive decoding and up to that of the strongest tree-speculative baseline. Under online agent workloads, it reduces TTFT and TPOT by up to and , respectively, over the strongest tree-speculative baseline.
Cite
@article{arxiv.2608.01651,
title = {Bole: Efficient Tree Speculation for Hybrid-Attention Language Models},
author = {Li Wang and Yi Su and Xiabao Wu and Chiran You and Yongchao Liu and Zhan Qiu and Juelu Zhang and Jiajun Zheng and Fangxin Liu and Jie Zhang and Chen Tian and Chengying Huan},
journal= {arXiv preprint arXiv:2608.01651},
year = {2026}
}
Comments
14 pages, 12 figures, 7 tables