Matrix-PIC: Harnessing Matrix Outer-product for High-Performance Particle-in-Cell Simulations
Abstract
Particle-in-Cell (PIC) simulations spend most of their execution time on particle--grid interactions, where fine-grained atomic updates become a major bottleneck on traditional many-core CPUs. Recent CPU architectures integrate specialized Matrix Processing Units (MPUs) that efficiently support matrix outer-product operations, offering new opportunities to overcome this limitation. Leveraging this architectural shift, this work focuses on redesigning the current deposition step of PIC simulations under a matrix-centric execution model. We present MatrixPIC, the first holistic co-design of the deposition kernel, data layout, and incremental particle sorting tailored to the hybrid MPU--VPU SIMD model on modern CPUs. MatrixPIC introduces: (i)~a block-matrix formulation of the current deposition algorithm that maps naturally to MPU outer-product primitives; (ii)~a hybrid execution pipeline that combines MPU-based high-density accumulation with VPU-based data preparation and control flow; and (iii)~an -amortized incremental sorter based on a gapped packed-memory array to preserve data locality for efficient MPU execution. Evaluated on a next-generation HPC platform, MatrixPIC achieves significant performance gains. In Laser-Wakefield Acceleration (LWFA) simulations, it delivers up to speedup in total runtime. For third-order deposition, the core kernel is accelerated by over the baseline and over the best hand-optimized VPU implementation. Moreover, MatrixPIC reaches of theoretical CPU peak performance, nearly higher than a highly optimized CUDA kernel on a data center GPU. These results demonstrate the effectiveness of matrix-oriented co-design for accelerating PIC simulations on emerging CPU architectures.
Cite
@article{arxiv.2601.08277,
title = {Matrix-PIC: Harnessing Matrix Outer-product for High-Performance Particle-in-Cell Simulations},
author = {Yizhuo Rao and Xingjian Cui and Jiabin Xie and Shangzhi Pang and Guangnan Feng and Jinhui Wei and Zhiguang Chen and Yutong Lu},
journal= {arXiv preprint arXiv:2601.08277},
year = {2026}
}
Comments
Accepted for publication at EuroSys 2026