English

NeCTAr: A Heterogeneous RISC-V SoC for Language Model Inference in Intel 16

Hardware Architecture 2025-03-20 v1

Abstract

This paper introduces NeCTAr (Near-Cache Transformer Accelerator), a 16nm heterogeneous multicore RISC-V SoC for sparse and dense machine learning kernels with both near-core and near-memory accelerators. A prototype chip runs at 400MHz at 0.85V and performs matrix-vector multiplications with 109 GOPs/W. The effectiveness of the design is demonstrated by running inference on a sparse language model, ReLU-Llama.

Keywords

Cite

@article{arxiv.2503.14708,
  title  = {NeCTAr: A Heterogeneous RISC-V SoC for Language Model Inference in Intel 16},
  author = {Viansa Schmulbach and Jason Kim and Ethan Gao and Lucy Revina and Nikhil Jha and Ethan Wu and Borivoje Nikolic},
  journal= {arXiv preprint arXiv:2503.14708},
  year   = {2025}
}
R2 v1 2026-06-28T22:25:57.508Z