ARM SVE大放异:面向Nvidia Grace的HPC应用性能与洞察
分布式、并行与集群计算
2025-05-15 v1
摘要
向量架构是提升计算吞吐量的关键。ARM提供的SVE作为下一代长度无关向量扩展,超越了传统固定长度SIMD。本工作为SVE在HPC领域的成熟度和就绪性提供了首个研究。我们选取ARM Grace处理器上的特定性能硬件事件和分析模型,构建新指标以量化SVE向量化减少执行指令数并提升性能加速的效果。我们进一步提出了结合向量长度与数据元素的适应版roofline模型,以识别潜在的性能瓶颈。最后,我们提出了一种用于分类SVE提升后性能的决策树。
关键词
引用
@article{arxiv.2505.09462,
title = {ARM SVE Unleashed: Performance and Insights Across HPC Applications on Nvidia Grace},
author = {Ruimin Shi and Gabin Schieffer and Maya Gokhale and Pei-Hung Lin and Hiren Patel and Ivy Peng},
journal= {arXiv preprint arXiv:2505.09462},
year = {2025}
}
备注
To be published in the 31st European Conference on Parallel and Distributed Processing(Euro-Par 2025)