English

31.1 A 14.08-to-135.69Token/s ReRAM-on-Logic Stacked Outlier-Free Large-Language-Model Accelerator with Block-Clustered Weight-Compression and Adaptive Parallel-Speculative-Decoding

Hardware Architecture 2026-05-12 v1

Abstract

This work presents a 55nm speculative decoding-based LLM accelerator with bumping-based face-to-face ReRAM-on-logic stacking technology. It features a local rotation unit for outlier-free low-bit quantization, a stacking-aware PNM architecture co-designed with blockwise vector quantization to reduce weight EMA overheads, and an adaptive parallel speculative decoding scheme with an out-of-order scheduler for high resource and bandwidth utilization. Our chip achieves 14.08-to-135.69token/s and 4.46-to-7.17x speedup over vanilla speculative decoding.

Keywords

Cite

@article{arxiv.2605.09375,
  title  = {31.1 A 14.08-to-135.69Token/s ReRAM-on-Logic Stacked Outlier-Free Large-Language-Model Accelerator with Block-Clustered Weight-Compression and Adaptive Parallel-Speculative-Decoding},
  author = {Pingcheng Dong and Yonghao Tan and Xuejiao Liu and Peng Luo and Yu Liu and Di Pang and Songchen Ma and Xijie Huang and Shih-Yang Liu and Dong Zhang and Zhichao Lu and Luhong Liang and Chi-Ying Tsui and Fengbin Tu and Liang Zhao and Kwang-Ting Cheng},
  journal= {arXiv preprint arXiv:2605.09375},
  year   = {2026}
}