English

AutoTSMM: An Auto-tuning Framework for Building High-Performance Tall-and-Skinny Matrix-Matrix Multiplication on CPUs

Distributed, Parallel, and Cluster Computing 2025-01-24 v2

Abstract

In recent years, general matrix-matrix multiplication with non-regular-shaped input matrices has been widely used in many applications like deep learning and has drawn more and more attention. However, conventional implementations are not suited for non-regular-shaped matrix-matrix multiplications, and few works focus on optimizing tall-and-skinny matrix-matrix multiplication on CPUs. This paper proposes an auto-tuning framework, AutoTSMM, to build high-performance tall-and-skinny matrix-matrix multiplication. AutoTSMM selects the optimal inner kernels in the install-time stage and generates an execution plan for the pre-pack tall-and-skinny matrix-matrix multiplication in the runtime stage. Experiments demonstrate that AutoTSMM achieves competitive performance comparing to state-of-the-art tall-and-skinny matrix-matrix multiplication. And, it outperforms all conventional matrix-matrix multiplication implementations.

Keywords

Cite

@article{arxiv.2208.08088,
  title  = {AutoTSMM: An Auto-tuning Framework for Building High-Performance Tall-and-Skinny Matrix-Matrix Multiplication on CPUs},
  author = {Chendi Li and Haipeng Jia and Hang Cao and Jianyu Yao and Boqian Shi and Chunyang Xiang and Jinbo Sun and Pengqi Lu and Yunquan Zhang},
  journal= {arXiv preprint arXiv:2208.08088},
  year   = {2025}
}

Comments

8 pages, 12 figures, published in IEEE ISPA 2021