PIP-NTT: Towards a Scalable Memory-Parallelized Accelerator for Iterative NTT in PQC
Abstract
The iterative forward and inverse number theoretic transform (NTT) is a key component in lattice-based post-quantum cryptography (PQC), typically implemented using Cooley-Tukey and Gentleman-Sande butterfly units. Existing iterative NTT accelerators often rely on ping-pong memory schemes and large memory blocks tied to the cyclotomic ring, which limits overall efficiency. To overcome this, we propose a memory-parallelization strategy using four smaller n/4-sized memories for ring size n, preserving the total memory footprint of conventional designs. We also introduce a multiplication-free rescaling architecture for the inverse NTT. Building on these innovations, we perform a comprehensive hardware-based design space exploration of unified Cooley-Tukey and Gentleman-Sande butterfly units, evaluating both coarse- and fine-grained pipelining strategies. The resulting optimized butterfly unit forms the core of our proposed pipelined and memory-parallelized NTT accelerator, "PIP-NTT". It integrates two such units alongside the memory-parallelization scheme to boost computational throughput under tight area constraints. Experimental results on FPGA platforms show that PIP-NTT achieves 2.67x and 1.48x higher efficiency in average Area-Time Product compared to the most area-optimized and high-speed NTT accelerators in the literature. The design is scalable across butterfly radices and adaptable to other PQC schemes, making it a versatile solution for future cryptographic hardware
Cite
@article{arxiv.2607.18533,
title = {PIP-NTT: Towards a Scalable Memory-Parallelized Accelerator for Iterative NTT in PQC},
author = {Malik Imran and Ayesha Khalid and Ciara Rafferty and Safiullah Khan and Muhammad Rashid and Maire O'Neill},
journal= {arXiv preprint arXiv:2607.18533},
year = {2026}
}
Comments
12 pages, 6 figures, 3 tables, Accepted in IEEE Transactions on Emerging Topics in Computing (TETC)