With Shared Microexponents, A Little Shifting Goes a Long Way
Machine Learning
2023-04-14 v2 Artificial Intelligence
Hardware Architecture
Abstract
This paper introduces Block Data Representations (BDR), a framework for exploring and evaluating a wide spectrum of narrow-precision formats for deep learning. It enables comparison of popular quantization standards, and through BDR, new formats based on shared microexponents (MX) are identified, which outperform other state-of-the-art quantization approaches, including narrow-precision floating-point and block floating-point. MX utilizes multiple levels of quantization scaling with ultra-fine scaling factors based on shared microexponents in the hardware. The effectiveness of MX is demonstrated on real-world models including large-scale generative pretraining and inferencing, and production-scale recommendation systems.
Cite
@article{arxiv.2302.08007,
title = {With Shared Microexponents, A Little Shifting Goes a Long Way},
author = {Bita Rouhani and Ritchie Zhao and Venmugil Elango and Rasoul Shafipour and Mathew Hall and Maral Mesmakhosroshahi and Ankit More and Levi Melnick and Maximilian Golub and Girish Varatkar and Lei Shao and Gaurav Kolhe and Dimitry Melts and Jasmine Klar and Renee L'Heureux and Matt Perry and Doug Burger and Eric Chung and Zhaoxia Deng and Sam Naghshineh and Jongsoo Park and Maxim Naumov},
journal= {arXiv preprint arXiv:2302.08007},
year = {2023}
}