Related papers: A computational system for lattice QCD with overla…
We simulate quenched QCD with the overlap Dirac operator. We work with the Wilson gauge action at beta=6 on an 18^3x64 lattice. We calculate quark propagators for a single source point and quark mass ranging from am_q=0.03 to 0.75. We…
We present here the most recent version of FermiQCD, a collection of C++ classes, functions and parallel algorithms for lattice QCD, based on Matrix Distributed Processing. FermiQCD allows fast development of parallel lattice applications…
We study the algorithmic optimization and performance tuning of the Lattice QCD clover-fermion solver for the K computer. We implement the L\"uscher's SAP preconditioner with sub-blocking in which the lattice block in a node is further…
The presence of GPU from different vendors demands the Lattice QCD codes to support multiple architectures. To this end, Open Computing Language (OpenCL) is one of the viable frameworks for writing a portable code. It is of interest to find…
Lattice QCD calculations require significant computational effort, with the dominant fraction of resources typically spent in the numerical inversion of the Dirac operator. One of the simplest methods to solve such large and sparse linear…
Quantum computers, with parallel computing and entanglement effects, excel in cryptography analysis and big data processing. However, they are not fully developed yet, and their performance needs further evaluation. Traditional computer…
Quantum error correction (QEC) and fault-tolerant (FT) mechanisms are essential for reliable quantum computing. However, QEC considerably increases the computation size up to four orders of magnitude. Moreover, FT implementation has…
I review recent machine trends and algorithmic developments for dynamical lattice QCD simulations with the HMC algorithm for Wilson-type fermions. The topics include the trend toward multi-core processors and general purpose GPU (GPGPU)…
We examine quenched chiral logarithms in lattice QCD with overlap Dirac quark. For 100 gauge configurations generated with the Wilson gauge action at $ \beta = 5.8 $ on the $ 8^3 \times 24 $ lattice, we compute quenched quark propagators…
We report on the status of the dynamical overlap QCD simulation project by the JLQCD collaboration. After completing two-flavor QCD simulation on a 16^3x32 lattice at lattice spacing a 0.12 fm, we started a series of runs with 2+1 flavors.…
We propose a novel general approach to locality of lattice composite fields, which in case of QCD involves locality in both quark and gauge degrees of freedom. The method is applied to gauge operators based on the overlap Dirac matrix…
The CP-PACS is a massively parallel computer dedicated for calculations in computational physics and will be in operation in the spring of 1996 at Center for Computational Physics, University of Tsukuba. In this article, we describe the…
Quantum low-density parity-check (qLDPC) codes can achieve high encoding rates and good code distance scaling, providing a promising route to low-overhead fault-tolerant quantum computing. However, the long-range connectivity required to…
Markov Chain Monte Carlo simulations of lattice Quantum Chromodynamics (QCD) are the only known tool to investigate non-perturbatively the theory of the strong interaction and are required to perform precision tests of the Standard Model of…
This paper describes a state-of-the-art parallel Lattice QCD Monte Carlo code for staggered fermions, purposely designed to be portable across different computer architectures, including GPUs and commodity CPUs. Portability is achieved…
We summarize our recent investigations of lattice QCD with dynamical overlap fermions. We sketch algorithmic issues and our approach to solving them. We show our measurement of the topological susceptibility. We describe a computation of…
The CP-PACS Project, which started in April 1992, is a five-year plan to develop a massively parallel computer for carrying out research in computational physics with primary emphasis on lattice QCD. This article describes the architectural…
We discuss the hardware design choices made in our 16K-node 0.8 Teraflops supercomputer project, a machine architecture optimized for full QCD calculations. The efficiency of the conjugate gradient algorithm in terms of balance of…
The QCDSP machine at Columbia University has grown to 2,048 nodes achieving a peak speed of 100 Gigaflops. Software for quenched and Hybrid Monte Carlo (HMC) evolution schemes has been developed for staggered fermions, with support for…
This study presents a benchmarking analysis of the Qualcomm Cloud AI 100 Ultra (QAic) accelerator for large language model (LLM) inference, evaluating its energy efficiency (throughput per watt), performance, and hardware scalability…