Related papers: Optimizing the domain wall fermion Dirac operator …
We review a number of topics related to block variable renormalisation group transformations of quantum fields on the lattice, and to the emerging perfect lattice actions. We first illustrate this procedure by considering scalar fields.…
Finite volume renormalization scheme is one of the most fascinating scheme for non-perturbative renormalization on lattice. By using the step scaling function one can follow running of renormalized quantities with reasonable cost. It has…
We study the algorithmic optimization and performance tuning of the Lattice QCD clover-fermion solver for the K computer. We implement the L\"uscher's SAP preconditioner with sub-blocking in which the lattice block in a node is further…
We propose a new architecture for distributed image compression from a group of distributed data sources. The work is motivated by practical needs of data-driven codec design, low power consumption, robustness, and data privacy. The…
With the upgrade of the RPCs [1]-[2] and the increase of its performances, the study and the optimization of the read-out panel is necessary in order to maintain the signal integrity and to reduce the intrinsic crosstalk. Through…
In this preliminary study, we examine the chiral properties of the parametrized Fixed-Point Dirac operator D^FP, see how to improve its chirality via the Overlap construction, measure the renormalized quark condensate Sigma and the…
Artificial monopoles have been engineered in various systems, yet there has been no systematic study of the singular vector potentials associated with the monopole field. We show that the Dirac string, the line singularity of the vector…
Over the past five years, graphics processing units (GPUs) have had a transformational effect on numerical lattice quantum chromodynamics (LQCD) calculations in nuclear and particle physics. While GPUs have been applied with great success…
The rapidly increasing number of cores available in multicore processors does not necessarily lead directly to a commensurate increase in performance: programs written in conventional languages, such as C, need careful restructuring,…
Developing scalable, fault-tolerant atomic quantum processors requires precise control over large arrays of optical beams. This remains a major challenge due to inherent imperfections in classical control hardware, such as inter-channel…
The supercomputing platforms available for high performance computing based research evolve at a great rate. However, this rapid development of novel technologies requires constant adaptations and optimizations of the existing codes for…
Resistive Random Access Memory (RRAM) and Phase Change Memory (PCM) devices have been popularly used as synapses in crossbar array based analog Neural Network (NN) circuit to achieve more energy and time efficient data classification…
This paper proposes distributed algorithms to solve robust convex optimization (RCO) when the constraints are affected by nonlinear uncertainty. We adopt a scenario approach by randomly sampling the uncertainty set. To facilitate the…
In this paper, we introduce a new coarse space algorithm, the "Discontinuous Coarse Space Robin Jump Minimizer" (DCS-RJMin), to be used in conjunction with one-level domain decomposition methods (DDM). This new algorithm makes use of…
We derive the effective action of the light fermion field of the domain-wall fermion, which is referred as q(x) and \bar q(x) by Furman and Shamir. The inverse of the effective Dirac operator turns out to be identical to the inverse of the…
OpenACC compilers allow one to use Graphics Processing Units without having to write explicit CUDA codes. Programs can be modified incrementally using OpenMP like directives which causes the compiler to generate CUDA kernels to be run on…
Variational optimization of neural-network representations of quantum states has been successfully applied to solve interacting fermionic problems. Despite rapid developments, significant scalability challenges arise when considering…
In this paper we would like to share our experience for transforming a parallel code for a Computational Fluid Dynamics (CFD) problem into a parallel version for the RedisDG workflow engine. This system is able to capture heterogeneous and…
Many scientific applications consist of large and computationally-intensive loops. Dynamic loop self-scheduling (DLS) techniques are used to parallelize and to balance the load during the execution of such applications. Load imbalance…
As the computing landscape evolves, system designers continue to explore design methodologies that leverage increased levels of heterogeneity to push performance within limited size, weight, power, and cost budgets. One such methodology is…