中文

ARM SVE大放异:面向Nvidia Grace的HPC应用性能与洞察

分布式、并行与集群计算 2025-05-15 v1

摘要

向量架构是提升计算吞吐量的关键。ARM提供的SVE作为下一代长度无关向量扩展,超越了传统固定长度SIMD。本工作为SVE在HPC领域的成熟度和就绪性提供了首个研究。我们选取ARM Grace处理器上的特定性能硬件事件和分析模型,构建新指标以量化SVE向量化减少执行指令数并提升性能加速的效果。我们进一步提出了结合向量长度与数据元素的适应版roofline模型,以识别潜在的性能瓶颈。最后,我们提出了一种用于分类SVE提升后性能的决策树。

关键词

引用

@article{arxiv.2505.09462,
  title  = {ARM SVE Unleashed: Performance and Insights Across HPC Applications on Nvidia Grace},
  author = {Ruimin Shi and Gabin Schieffer and Maya Gokhale and Pei-Hung Lin and Hiren Patel and Ivy Peng},
  journal= {arXiv preprint arXiv:2505.09462},
  year   = {2025}
}

备注

To be published in the 31st European Conference on Parallel and Distributed Processing(Euro-Par 2025)