中文

面向嵌入式系统高效机器学习的编译与优化

机器学习 2022-08-29 v2 硬件体系结构

摘要

深度神经网络(DNNs)已在多种机器学习(ML)应用中取得巨大成功,在计算机视觉、自然语言处理及虚拟现实等领域提供高质量推理方案。然而,基于 DNN 的 ML 应用也带来了大幅增长的计算与存储需求,这对计算/存储资源有限、功耗预算紧张且外形尺寸小的嵌入式系统尤为棘手。挑战还来自多样的应用特定需求,包括实时响应、高吞吐性能与可靠推理精度。为应对这些挑战,我们引入一系列有效设计方法,包括高效的 ML 模型设计、定制化硬件加速器设计,以及软硬件协同设计策略,以在嵌入式系统上实现高效的 ML 应用。

关键词

引用

@article{arxiv.2206.03326,
  title  = {Compilation and Optimizations for Efficient Machine Learning on Embedded Systems},
  author = {Xiaofan Zhang and Yao Chen and Cong Hao and Sitao Huang and Yuhong Li and Deming Chen},
  journal= {arXiv preprint arXiv:2206.03326},
  year   = {2022}
}

备注

This article will appear as a book chapter in a new book: Embedded Machine Learning for Cyber-Physical, IoT, and Edge Computing, Springer Nature