面向 MLPerf Tiny 基准测试的开源 FPGA-机器学习协同设计
机器学习
2022-06-24 v1 硬件体系结构
摘要
我们展示了在 FPGA(现场可编程门阵列)平台上针对 MLPerf Tiny 推理基准测试的开发经验与近期成果。我们使用了开源的 hls4ml 与 FINN 工作流,其旨在使 FPGA 上优化神经网络的 AI-硬件协同设计民主化。我们展示了关键词 spotting、异常检测与图像分类基准任务的设计与实现流程。所得的硬件实现为量化、可配置、空间数据流架构,专为速度与效率定制,并引入了作为本工作一部分而开发的新的通用优化与通用工作流。完整工作流从量化感知训练到 FPGA 实现均有呈现。相关方案部署于片上系统(Pynq-Z2)与纯 FPGA(Arty A7-100T)平台。所得提交方案的推理延迟低至 20 s,能耗低至每次推理 30 J。我们展示了异构硬件平台上新兴的 ML 基准测试如何催化协作以及新技术与更易用工具的发展。
引用
@article{arxiv.2206.11791,
title = {Open-source FPGA-ML codesign for the MLPerf Tiny Benchmark},
author = {Hendrik Borras and Giuseppe Di Guglielmo and Javier Duarte and Nicolò Ghielmetti and Ben Hawks and Scott Hauck and Shih-Chieh Hsu and Ryan Kastner and Jason Liang and Andres Meza and Jules Muhizi and Tai Nguyen and Rushil Roy and Nhan Tran and Yaman Umuroglu and Olivia Weng and Aidan Yokuda and Michaela Blott},
journal= {arXiv preprint arXiv:2206.11791},
year = {2022}
}
备注
15 pages, 7 figures, Contribution to 3rd Workshop on Benchmarking Machine Learning Workloads on Emerging Hardware (MLBench) at 5th Conference on Machine Learning and Systems (MLSys)