English

Training and inference of large language models using 8-bit floating point

Machine Learning 2023-10-02 v1 Hardware Architecture Computation and Language Emerging Technologies Performance

Abstract

FP8 formats are gaining popularity to boost the computational efficiency for training and inference of large deep learning models. Their main challenge is that a careful choice of scaling is needed to prevent degradation due to the reduced dynamic range compared to higher-precision formats. Although there exists ample literature about selecting such scalings for INT formats, this critical aspect has yet to be addressed for FP8. This paper presents a methodology to select the scalings for FP8 linear layers, based on dynamically updating per-tensor scales for the weights, gradients and activations. We apply this methodology to train and validate large language models of the type of GPT and Llama 2 using FP8, for model sizes ranging from 111M to 70B. To facilitate the understanding of the FP8 dynamics, our results are accompanied by plots of the per-tensor scale distribution for weights, activations and gradients during both training and inference.

Keywords

Cite

@article{arxiv.2309.17224,
  title  = {Training and inference of large language models using 8-bit floating point},
  author = {Sergio P. Perez and Yan Zhang and James Briggs and Charlie Blake and Josh Levy-Kramer and Paul Balanca and Carlo Luschi and Stephen Barlow and Andrew William Fitzgibbon},
  journal= {arXiv preprint arXiv:2309.17224},
  year   = {2023}
}
R2 v1 2026-06-28T12:36:04.739Z