Related papers: A performance evaluation of CCS QCD Benchmark on t…
Graphical Processing Units (GPUs) are more and more frequently used for lattice QCD calculations. Lattice studies often require computing the quark propagators for several masses. These systems can be solved using multi-shift inverters but…
In these proceedings we address the computation of quark-line disconnected diagrams in lattice QCD. The evaluation of these diagrams is required for many phenomenologically interesting observables, but suffers from large statistical errors…
We present a comparison of a number of iterative solvers of linear systems of equations for obtaining the fermion propagator in lattice QCD. In particular, we consider chirally invariant overlap and chirally improved Wilson (maximally)…
PLQCD is a stand-alone software library developed under PRACE for lattice QCD. It provides an implementation of the Dirac operator for Wilson type fermions and few efficient linear solvers. The library is optimized for multi-core machines…
We outline the essential features of a Linux PC cluster which is now being developed at National Taiwan University, and discuss how to optimize its hardware and software for lattice QCD with overlap Dirac quarks. At present, the cluster…
Deploying new supercomputers requires testing and evaluation via application codes. Portable, user-friendly tools enable evaluation, and the Multicomponent Flow Code (MFC), a computational fluid dynamics (CFD) code, addresses this need. MFC…
We investigate the use of half-precision floating-point numbers (FP16) in mixed-precision linear solvers for lattice QCD simulations. Since the emergence of GPUs for general-purpose, mixed-precision algorithms that combine single-precision…
The computational effort in the calculation of Wilson fermion quark propagators in Lattice Quantum Chromodynamics can be considerably reduced by exploiting the Wilson fermion matrix structure in inversion algorithms based on the…
We explore the possibility of computing fermionic correlators on the lattice by combining a domain decomposition with a multi-level integration scheme. The quark propagator is expanded in series of terms with a well defined hierarchical…
Performance of distributed data center applications can be improved through use of FPGA-based SmartNICs, which provide additional functionality and enable higher bandwidth communication. Until lately, however, the lack of a simple approach…
First results of a recently started simulation of full QCD with two flavours of sea-quarks at a coupling of $\beta = 5.6$ on a $16^3 \times 32$ lattice are presented. Emphasis is laid on the statistical significance that can be achieved by…
We compare different conjugate gradient -- like matrix inversion methods (CG, BiCGstab1 and BiCGstab2) employing for this purpose the compact lattice quantum electrodynamics (QED) with Wilson fermions. The main goals of this investigation…
In the push for exascale computing, energy efficiency is of utmost concern. System architectures often adopt accelerators to hasten application execution at the cost of power. The Intel Xeon Phi co-processor is unique accelerator that…
We present a new exact algorithm for estimating all elements of the quark propagator. The advantage of the method is that the exact all-to-all propagator is reproduced in a large but finite number of inversions. The efficacy of the…
We simulate quenched QCD with the overlap Dirac operator. We work with the Wilson gauge action at beta=6 on an 18^3x64 lattice. We calculate quark propagators for a single source point and quark mass ranging from am_q=0.03 to 0.75. We…
Modern OpenMP threading techniques are used to convert the MPI-only Hartree-Fock code in the GAMESS program to a hybrid MPI/OpenMP algorithm. Two separate implementations that differ by the sharing or replication of key data structures…
We discuss the implementation of a Sheikholeslami-Wohlert term for simulations of lattice QCD with dynamical Wilson fermions as required by Symanzik's improvement program. We show that for the Hybrid Monte Carlo or Kramers equation…
Low bit-width Quantized Neural Networks (QNNs) enable deployment of complex machine learning models on constrained devices such as microcontrollers (MCUs) by reducing their memory footprint. Fine-grained asymmetric quantization (i.e.,…
In recent years the computational capacity of single Field Programmable Gate Arrays (FPGA) devices as well as their versatility has increased significantly. Adding to that the High Level Synthesis frameworks allowing to program such…
We investigate the computational efficiency of two stochastic based alternatives to the Sequential Propagator Method used in Lattice QCD calculations of heavy-light semileptonic form factors. In the first method, we replace the sequential…