English

A Study of Quantisation-aware Training on Time Series Transformer Models for Resource-constrained FPGAs

Machine Learning 2023-10-05 v1 Artificial Intelligence Hardware Architecture

Abstract

This study explores the quantisation-aware training (QAT) on time series Transformer models. We propose a novel adaptive quantisation scheme that dynamically selects between symmetric and asymmetric schemes during the QAT phase. Our approach demonstrates that matching the quantisation scheme to the real data distribution can reduce computational overhead while maintaining acceptable precision. Moreover, our approach is robust when applied to real-world data and mixed-precision quantisation, where most objects are quantised to 4 bits. Our findings inform model quantisation and deployment decisions while providing a foundation for advancing quantisation techniques.

Keywords

Cite

@article{arxiv.2310.02654,
  title  = {A Study of Quantisation-aware Training on Time Series Transformer Models for Resource-constrained FPGAs},
  author = {Tianheng Ling and Chao Qian and Lukas Einhaus and Gregor Schiele},
  journal= {arXiv preprint arXiv:2310.02654},
  year   = {2023}
}

Comments

12 pages, 1 figure