Intel Xeon Phi和NVIDIA GPU上的HISQ反演器
分布式、并行与集群计算
2014-09-05 v1 高能物理 - 格点
摘要
格点QCD模拟的运行时间主要由一个小内核决定,该内核计算向量与稀疏矩阵(称为“Dslash”算子)的乘积。因此,该内核经常针对各种HPC架构进行优化。本文比较了Intel Xeon Phi与当前基于Kepler的NVIDIA Tesla GPU在运行共轭梯度求解器时的性能。通过同时反转多个向量向加速器暴露更多并行性,我们在两种架构上都获得了250 GFlop/s的性能。这使反演性能提高了一倍以上。我们简要概述了两种架构,讨论了实现的一些细节以及获得该性能所需的工作量。
引用
@article{arxiv.1409.1510,
title = {HISQ inverter on Intel Xeon Phi and NVIDIA GPUs},
author = {O. Kaczmarek and C. Schmidt and P. Steinbrecher and Swagato Mukherjee and M. Wagner},
journal= {arXiv preprint arXiv:1409.1510},
year = {2014}
}
备注
7 pages, proceedings, presented at the 32nd International Symposium on Lattice Field Theory (Lattice 2014), June 23 to June 28, 2014, New York, USA