Related papers: Lattice QCD Applications on QPACE
In this report, I describe the design and implementation of an inexpensive, eight node, 32 core, cluster of raspberry pi single board computers, as well as the performance of this cluster on two computational tasks, one that requires…
The prospects of quantum computing have driven efforts to realize fully functional quantum processing units (QPUs). Recent success in developing proof-of-principle QPUs has prompted the question of how to integrate these emerging processors…
We summarize the status of lattice QCD ensemble generation efforts and their data management characteristics. Namely, these proceedings combine the contributions to a dedicated parallel session during the 41st International Symposium on…
We present full accounts of a method to extract nucleon-nucleon (NN) potentials from the Bethe-Salpter amplitude in lattice QCD. The method is applied to two nucleons on the lattice with quenched QCD simulations. By disentangling the mixing…
We report on the status of the dynamical overlap QCD simulation project by the JLQCD collaboration. After completing two-flavor QCD simulation on a 16^3x32 lattice at lattice spacing a 0.12 fm, we started a series of runs with 2+1 flavors.…
Quantum Data Centers (QDCs) are needed to support large-scale quantum processing for both academic and commercial applications. While large-scale quantum computers are constrained by technological and financial barriers, a modular approach…
What if you could piece together your own custom biometrics and AI analysis system, a bit like LEGO blocks? We aim to bring that technology to field operators in the field who require flexible, high-performance edge AI system that can be…
Systolic arrays and shared-L1-memory manycore clusters are commonly used architectural paradigms that offer different trade-offs to accelerate parallel workloads. While the first excel with regular dataflow at the cost of rigid…
One of the barriers to the adoption of parallel computing is the inherent complexity of its programming. The Open Multi-Processing (OpenMP) Application Programming Interface (API) facilitates such implementations, providing high abstraction…
Over the most recent years, quantized graph neural network (QGNN) attracts lots of research and industry attention due to its high robustness and low computation and memory overhead. Unfortunately, the performance gains of QGNN have never…
iPIC3D is a widely used massively parallel Particle-in-Cell code for the simulation of space plasmas. However, its current implementation does not support execution on multiple GPUs. In this paper, we describe the porting of iPIC3D particle…
To address the growing needs for scalable High Performance Computing (HPC) and Quantum Computing (QC) integration, we present our HPC-QC full stack framework and its hybrid workload development capability with modular…
We present a fault-tolerant universal quantum computing architecture based on a code concatenation of biased-noise qubits and the parity architecture. The parity architecture can be understood as an LDPC code tailored specifically to obtain…
Though CNNs are highly parallel workloads, in the absence of efficient on-chip memory reuse techniques, an accelerator for them quickly becomes memory bound. In this paper, we propose a CNN accelerator design for inference that is able to…
It is well-known that molecular dynamics integrators, which are used for lattice quantum chromodynamics (QCD), suffer from instabilities and possess a rather low order of the accuracy. Hence, it is highly desirable to construct a new class…
The lattice technique of studying the strong interaction of matter is used to obtain predictions of the hadronic spectrum. These simulations were performed by the UKQCD collaboration using full (unquenched) QCD. Details of the results, a…
The emerging field of quantum resource estimation is aimed at providing estimates of the hardware requirements (`quantum resources') needed to execute a useful, fault-tolerant quantum computation. Given that quantum computers are intended…
The limited number of qubits per chip remains a critical bottleneck in quantum computing, motivating the use of distributed architectures that interconnect multiple quantum processing units (QPUs). However, executing quantum algorithms…
The growing demand for efficient, high-performance processing in machine learning (ML) and image processing has made hardware accelerators, such as GPUs and Data Streaming Accelerators (DSAs), increasingly essential. These accelerators…
We present a footprint study for the scaling of modular quantum error correction (QEC) protocols designed for triangular color codes, including a lattice-surgery-based logical teleportation gadget, and compare the performance of various…