Intel Xeon Phi 架构上的 Staggered Dslash 性能
高能物理 - 格点
2014-11-11 v1 计算物理
摘要
共轭梯度(CG)算法是带有 staggered 夸克的格点计算中最基本且最耗时的部分之一。我们在 Intel Xeon Phi(也称为多集成核心,MIC 架构)上测试了 CG 和 dslash(CG 算法中的关键步骤)的性能。我们尝试了使用 MPI、OpenMP 和向量处理单元(VPUs)的不同并行化策略。
引用
@article{arxiv.1411.2087,
title = {Staggered Dslash Performance on Intel Xeon Phi Architecture},
author = {Ruizi Li and Steven Gottlieb},
journal= {arXiv preprint arXiv:1411.2087},
year = {2014}
}
备注
7 Pages, 1 figure, contribution to the 32nd International Symposium on Lattice Field Theory (Lattice 2014), 23-28 June 2014, Columbia University, New York, NY, USA