English

CIMERA: Compute-in-Interconnect and Memory with Reconfigurable Precision for LLM Inference

Hardware Architecture 2026-07-15 v1

Abstract

LLM impose significant computational and memory demands, creating challenges for energy-efficient inference across platforms ranging from data centers to power-constrained edge devices. Weight precision plays a critical role in balancing inference accuracy, throughput, and energy consumption, while modern LLM workloads exhibit pronounced heterogeneity and tolerance that favors adaptive precision execution. This paper presents CIMERA, a reconfigurable-precision LLM inference accelerator that integrates compute-in-interconnect and memory to mitigate the memory wall and enable precision-aware execution. Compared to Nvidia H100, CIMERA delivers up to 25×25\times and 10×10\times higher energy efficiency for 1B and 13B models, respectively.

Cite

@article{arxiv.2607.13649,
  title  = {CIMERA: Compute-in-Interconnect and Memory with Reconfigurable Precision for LLM Inference},
  author = {Yue Jiet Chong and Yimin Wang and Wei Zhang and Xuanyao Fong},
  journal= {arXiv preprint arXiv:2607.13649},
  year   = {2026}
}

Comments

Accepted to 2026 IEEE 8th International Conference on Artificial Intelligence Circuits and Systems (AICAS'26)