Intel Xeon Phi 与 NVIDIA GPU 上的共轭梯度求解器
计算物理
2014-11-18 v1 数学软件
高能物理 - 格点
摘要
格点量子色动力学模拟通常将大部分运行时间花费在费米子矩阵的求逆上。因此,这部分经常针对各种高性能计算架构进行优化。在此,我们比较了运行共轭梯度求解器时 Intel Xeon Phi 与当前基于 Kepler 架构的 NVIDIA Tesla GPU 的性能。通过同时求逆多个向量以向加速器暴露更多并行性,我们在两种架构上均获得了超过 300 GFlop/s 的性能。这使得求逆性能提高了一倍以上。我们还简要概述了 Knights Corner 架构,讨论了实现的一些细节以及为达到所实现性能所需的工作量。
引用
@article{arxiv.1411.4439,
title = {Conjugate gradient solvers on Intel Xeon Phi and NVIDIA GPUs},
author = {O. Kaczmarek and C. Schmidt and P. Steinbrecher and M. Wagner},
journal= {arXiv preprint arXiv:1411.4439},
year = {2014}
}
备注
7 pages, proceedings, presented at 'GPU Computing in High Energy Physics', September 10-12, 2014, Pisa, Italy