嵌入式 FPGA 中加速混合极低位宽神经网络的设计流程
分布式、并行与集群计算
2018-10-29 v2 计算机视觉与模式识别
机器学习
摘要
具有低延迟和低能耗的神经网络加速器对于边缘计算十分理想。为构建此类加速器,我们提出了一种在采用混合量化方案的嵌入式 FPGA 中加速极低位宽神经网络(ELB-NN)的设计流程。该流程涵盖网络训练与基于 FPGA 的网络部署,有助于设计空间探索并简化网络精度与计算效率之间的权衡。使用此流程可帮助硬件设计人员在严格的资源和功耗约束下于边缘设备中交付网络加速器。我们通过支持神经网络内的混合 ELB 设置来展示所提出的流程。结果表明,我们的设计可提供峰值达 10.3 TOPS 的极高性能,并在使用嵌入式 FPGA 以低于 5W 功耗运行大规模神经网络时,每秒每瓦分类多达 325.3 张图像。据我们所知,与迄今文献中报道的 GPU 或其他 FPGA 实现相比,它是最节能的解决方案。
引用
@article{arxiv.1808.04311,
title = {Design Flow of Accelerating Hybrid Extremely Low Bit-width Neural Network in Embedded FPGA},
author = {Junsong Wang and Qiuwen Lou and Xiaofan Zhang and Chao Zhu and Yonghua Lin and Deming Chen},
journal= {arXiv preprint arXiv:1808.04311},
year = {2018}
}
备注
Accepted by International Conference on Field-Programmable Logic and Applications (FPL'2018)