Related papers: Optimizing the domain wall fermion Dirac operator …
The dominant cost of most lattice QCD simulations is the inversion of the Dirac operator required to calculate the force term in the RHMC update. One way to improve this situation is to use multiple pseudofermions, which reduces the size…
Stream computation is one of the approaches suitable for FPGA-based custom computing due to its high throughput capability brought by pipelining with regular memory access. To increase performance of iterative stream computation, we can…
We construct new Ginsparg-Wilson fermions for QCD by inserting an approximately chiral Dirac operator - which involves ingredients of a perfect action - into the overlap formula. This accelerates the convergence of the overlap Dirac…
We present an exact pseudofermion action for hybrid Monte Carlo simulation (HMC) of one-flavor domain-wall fermion (DWF), with the effective 4-dimensional Dirac operator equal to the optimal rational approximation of the overlap-Dirac…
Simulating lattice QCD with chiral fermions and indeed using Domain Wall Fermions continues to be challenging project however large are concurrent computers. One obvious bottleneck is the slow pace of prototyping using the low level coding…
Computational Fluid Dynamics (CFD) simulations are often constrained by the memory-bound nature of sparse matrix-vector operations, which eventually limits performance on modern high-performance computing (HPC) systems. This work introduces…
Optical computing is considered a promising solution for the growing demand for parallel computing in various cutting-edge fields, requiring high integration and high speed computational capacity. In this paper, we propose a novel optical…
In lattice QCD computations a substantial amount of work is spent in solving linear systems arising in Wilson's discretization of the Dirac equations. We show first numerical results of the extension of the two-level DD-\alpha AMG method to…
Stochastic Computing (SC) is a computing paradigm that allows for the low-cost and low-power computation of various arithmetic operations using stochastic bit streams and digital logic. In contrast to conventional representation schemes…
The challenges associated with effectively programming FPGAs have been a major blocker in popularising reconfigurable architectures for HPC workloads. However new compiler technologies, such as MLIR, are providing new capabilities which…
Computing demands for large scientific experiments, such as the CMS experiment at the CERN LHC, will increase dramatically in the next decades. To complement the future performance increases of software running on central processing units…
We report on the first successful QCD multigrid algorithm which demonstrates constant convergence rates independent of quark mass and lattice volume for the Wilson Dirac operator. The new ingredient is the adaptive method for constructing…
Fault tolerance is a property which needs deeper consideration when dealing with streaming jobs requiring high levels of availability and low-latency processing even in case of failures where Quality-of-Service constraints must be adhered…
We compute the chiral condensate in 2+1-flavor QCD through the spectrum of low-lying eigenmodes of Dirac operator. The number of eigenvalues of the Dirac operator is evaluated using a stochastic method with an eigenvalue filtering technique…
Applying domain decomposition to the lattice Dirac operator and the associated quark propagator, we arrive at expressions which, with the proper insertion of random sources therein, can provide improvement to the estimation of the…
Massively parallel accelerators such as GPGPUs, manycores and FPGAs represent a powerful and affordable tool for scientists who look to speed up simulations of complex systems. However, porting code to such devices requires a detailed…
During the last two years, the goal of many researchers has been to squeeze the last bit of performance out of HPC system for AI tasks. Often this discussion is held in the context of how fast ResNet50 can be trained. Unfortunately,…
Domain-Specific Languages (DSLs) improve programmers productivity by decoupling problem descriptions from algorithmic implementations. However, DSLs for High-Performance Computing (HPC) have two additional critical requirements: performance…
We present the domain-wall fermion operator which is reflection symmetric in the fifth dimension, with the approximate sign function $ S(H) $ of the effective 4-dimensional Dirac operator satisfying the bound $ |1-S(\lambda)| \le 2 d_Z $…
Applying domain decomposition to the lattice Dirac operator and the associated quark propagator, we arrive at expressions which, with the proper insertion of random sources therein, can provide improvement to the estimation of the…