具有偏斜流水线的 systolic 阵列中的精简精度浮点运算
硬件体系结构
2023-09-11 v3
摘要
深度学习核在硬件上的加速依赖于在 systolic 阵列(SA)上高效执行的矩阵乘法。为有效权衡深度学习训练/推理质量与硬件成本,SA 加速器采用精简精度浮点(FP)运算。在本工作中,我们论证了需要新的流水线组织来降低精简精度 FP 运算单元的延迟并提升能效,以应对 SA 结构所施加的链式乘加运算。所提出的偏斜流水线设计重组了 FP 乘加单元的流水化操作,为指数逻辑启用新的转发路径,从而允许连续 PE 的流水线级并行执行。结果,SA 内矩阵乘法操作的延迟以极小硬件代价显著降低,从而为所考察的先进 CNN 带来 8% 和 11% 的能耗降低。
引用
@article{arxiv.2304.01668,
title = {Reduced-Precision Floating-Point Arithmetic in Systolic Arrays with Skewed Pipelines},
author = {D. Filippas and C. Peltekis and G. Dimitrakopoulos and C. Nicopoulos},
journal= {arXiv preprint arXiv:2304.01668},
year = {2023}
}
备注
IEEE International Conference on Artificial Intelligence Circuits and Systems (AICAS) 2023