中文

Intel Xeon Phi和NVIDIA GPU上的HISQ反演器

分布式、并行与集群计算 2014-09-05 v1 高能物理 - 格点

摘要

格点QCD模拟的运行时间主要由一个小内核决定,该内核计算向量与稀疏矩阵(称为“Dslash”算子)的乘积。因此,该内核经常针对各种HPC架构进行优化。本文比较了Intel Xeon Phi与当前基于Kepler的NVIDIA Tesla GPU在运行共轭梯度求解器时的性能。通过同时反转多个向量向加速器暴露更多并行性,我们在两种架构上都获得了250 GFlop/s的性能。这使反演性能提高了一倍以上。我们简要概述了两种架构,讨论了实现的一些细节以及获得该性能所需的工作量。

关键词

引用

@article{arxiv.1409.1510,
  title  = {HISQ inverter on Intel Xeon Phi and NVIDIA GPUs},
  author = {O. Kaczmarek and C. Schmidt and P. Steinbrecher and Swagato Mukherjee and M. Wagner},
  journal= {arXiv preprint arXiv:1409.1510},
  year   = {2014}
}

备注

7 pages, proceedings, presented at the 32nd International Symposium on Lattice Field Theory (Lattice 2014), June 23 to June 28, 2014, New York, USA