CIMERA: Compute-in-Interconnect and Memory with Reconfigurable Precision for LLM Inference
Abstract
LLM impose significant computational and memory demands, creating challenges for energy-efficient inference across platforms ranging from data centers to power-constrained edge devices. Weight precision plays a critical role in balancing inference accuracy, throughput, and energy consumption, while modern LLM workloads exhibit pronounced heterogeneity and tolerance that favors adaptive precision execution. This paper presents CIMERA, a reconfigurable-precision LLM inference accelerator that integrates compute-in-interconnect and memory to mitigate the memory wall and enable precision-aware execution. Compared to Nvidia H100, CIMERA delivers up to and higher energy efficiency for 1B and 13B models, respectively.
Cite
@article{arxiv.2607.13649,
title = {CIMERA: Compute-in-Interconnect and Memory with Reconfigurable Precision for LLM Inference},
author = {Yue Jiet Chong and Yimin Wang and Wei Zhang and Xuanyao Fong},
journal= {arXiv preprint arXiv:2607.13649},
year = {2026}
}
Comments
Accepted to 2026 IEEE 8th International Conference on Artificial Intelligence Circuits and Systems (AICAS'26)