Related papers: Optimizing the domain wall fermion Dirac operator …
Data stream processing systems (DSPSs) enable users to express and run stream applications to continuously process data streams. To achieve real-time data analytics, recent researches keep focusing on optimizing the system latency and…
An alternative to commonly used domain wall fermions is presented. Some rigorous bounds on the condition number of the associated linear problem are derived. On the basis of these bounds and some experimentation it is argued that domain…
An overview is given of the QCDOC architecture, a massively parallel and highly scalable computer optimized for lattice QCD using system-on-a-chip technology. The heart of a single node is the PowerPC-based QCDOC ASIC, developed in…
For a multi-cell, multi-user, cellular network downlink sum-rate maximization through power allocation is a nonconvex and NP-hard optimization problem. In this paper, we present an effective approach to solving this problem through single-…
Sparse-dense linear algebra is crucial in many domains, but challenging to handle efficiently on CPUs, GPUs, and accelerators alike; multiplications with sparse formats like CSR and CSF require indirect memory lookups. In this work, we…
New hardware architectures open up immense opportunities for supercomputer simulations. However, programming techniques for different architectures vary significantly, which leads to the necessity of developing and supporting multiple code…
Performance evaluations on the deterministic algorithms for 6-D problems are rarely found in literatures except some recent advances in the Vlasov and Boltzmann community [Dimarco et al. (2018), Kormann et al. (2019)], due to the extremely…
Designing stable cluster synchronization patterns is a fundamental challenge in nonlinear dynamics of networks with great relevance to understanding neuronal and brain dynamics. So far, cluster synchronization has been studied exclusively…
We present our implementation of the RHMC algorithm for staggered fermions on Graphics Processing Units using the NVIDIA CUDA programming language. While previous studies exclusively deal with the Dirac matrix inversion problem, our code…
No area of computing is hungrier for performance than High Performance Computing (HPC), the demands of which continue to be a major driver for processor performance and adoption of accelerators, and also advances in memory, storage, and…
Scientific applications often contain large, computationally-intensive, and irregular parallel loops or tasks that exhibit stochastic characteristics. Applications may suffer from load imbalance during their execution on high-performance…
Dirac particle dynamics is encoded as a unitary path summation rule and implemented on a qubit array, where the qubit array represents both spacetime and the fermions contained therein. The unitary path summation rule gives a quantum…
Traditional Digital Signal Processing ( DSP ) compilers work at low level ( C-level / assembly level ) and hence lose much of the optimization opportunities present at high-level ( domain-level ). The emerging multi-level compiler…
Efficient algorithms for the solution of partial differential equations on parallel computers are often based on domain decomposition methods. Schwarz preconditioners combined with standard Krylov space solvers are widely used in this…
We describe a way to optimize the chiral behavior of Wilson-type lattice fermion actions by studying the low energy real eigenmodes of the Dirac operator. We find a candidate action, the clover action with fat links with a tuned clover…
The extreme computational costs of calculating the sign of the Wilson matrix within the overlap operator have so far prevented four dimensional dynamical overlap simulations on realistic lattice sizes, because the computational power…
Dynamic scaling is critical to stream processing engines, as their long-running nature demands adaptive resource management. Existing scaling approaches easily cause performance degradation due to coarse-grained synchronization and…
Deep Reinforcement Learning (DRL) underlies in a simulated environment and optimizes objective goals. By extending the conventional interaction scheme, this paper proffers gym-ds3, a scalable and reproducible open environment tailored for a…
In a common approach for parallel processing applied to simulations of many-particle systems with short-ranged interactions and uniform density, the simulation cell is partitioned into domains of equal shape and size, each of which is…
We present a study of the Dirac eigenvalue spectrum near the region of the QCD phase transition. This study makes use of a sequence of ensembles with temperatures from 150 MeV to 200 MeV generated with $2 + 1$ flavors of dynamical domain…