English

Low Latency Transformer Inference on FPGAs for Physics Applications with hls4ml

Machine Learning 2024-09-10 v1

Abstract

This study presents an efficient implementation of transformer architectures in Field-Programmable Gate Arrays(FPGAs) using hls4ml. We demonstrate the strategy for implementing the multi-head attention, softmax, and normalization layer and evaluate three distinct models. Their deployment on VU13P FPGA chip achieved latency less than 2us, demonstrating the potential for real-time applications. HLS4ML compatibility with any TensorFlow-built transformer model further enhances the scalability and applicability of this work. Index Terms: FPGAs, machine learning, transformers, high energy physics, LIGO

Keywords

Cite

@article{arxiv.2409.05207,
  title  = {Low Latency Transformer Inference on FPGAs for Physics Applications with hls4ml},
  author = {Zhixing Jiang and Dennis Yin and Yihui Chen and Elham E Khoda and Scott Hauck and Shih-Chieh Hsu and Ekaterina Govorkova and Philip Harris and Vladimir Loncar and Eric A. Moreno},
  journal= {arXiv preprint arXiv:2409.05207},
  year   = {2024}
}
R2 v1 2026-06-28T18:37:54.192Z