Related papers: Two-link Staggered Quark Smearing in QUDA
The Linear Combination of Unitaries (LCU) method is a powerful scheme for the block encoding of operators but suffers from high overheads. In this work, we discuss the parallelisation of LCU and in particular the SELECT subroutine of LCU…
We present a high statistics study of the light hadron spectrum and quark masses in QCD with two flavors of dynamical quarks. Numerical simulations are carried out using the plaquette gauge action and the O(a)-improved Wilson quark action…
The alternating direction method of multipliers (ADMM) is a powerful operator splitting technique for solving structured convex optimization problems. Due to its relatively low per-iteration computational cost and ability to exploit…
With the increase in the amount of data and the expansion of model scale, distributed parallel training becomes an important and successful technique to address the optimization challenges. Nevertheless, although distributed stochastic…
Neural network training entails heavy computation with obvious bottlenecks. The Compute Unified Device Architecture (CUDA) programming model allows us to accelerate computation by passing the processing workload from the CPU to the graphics…
We present preliminary results from exploring the phase diagram of finite temperature QCD with three degenerate flavors and with two light flavors and the mass of the third held approximately at the strange quark mass. We use an order…
This paper describes the application of the code generated by the CAMPARY software to accelerate the solving of linear systems in the least squares sense on Graphics Processing Units (GPUs), in double double, quad double, and octo double…
Scanning probe microscopy with multi-qubit sensors offers the potential to improve imaging speed and measure previously inaccessible quantities, such as two-point correlations. We develop a multiplexed quantum sensing approach with scanning…
We present simulation details and results for the light hadron spectrum in N f = 2 + 1 lattice QCD with the nonperturbatively O(a)-improved Wilson quark action and the Iwasaki gauge action. Simulations are carried out at a lattice spacing…
We have extended our previous study of the lattice QCD spectrum with 2 flavors of staggered dynamical quarks at $6/g^2=5.6$ and $am_q=0.025$ and 0.01 to larger lattices, with better statistics and with additional sources for the…
This paper presents an experimental evaluation of parallel-in-time Kalman filters and smoothers using graphics processing units (GPUs). In particular, the paper evaluates different all-prefix-sum algorithms, that is, parallel scan…
Results of porting parts of the Lattice Quantum Chromodynamics code to modern FPGA devices are presented. A single-node, double precision implementation of the Conjugate Gradient algorithm is used to invert numerically the Dirac-Wilson…
Hybrid quantum-HPC algorithms advance research by delegating complex tasks to quantum processors and using HPC systems to orchestrate workflows and complementary computations. Sample-based quantum diagonalization (SQD) is a hybrid…
Quantum computing is emerging as an important (but radical) technology that might take us beyond Moore's law for certain applications. Today, in parallel with improving quantum computers, computer scientists are relying heavily on quantum…
We introduce a new class of actions for staggered quarks in lattice QCD which significantly reduce flavour symmetry violations in the pion mass spectrum. An action introduced by the MILC collaboration for the same purpose is seen to be a…
Taste symmetry violations in staggered fermion formulations correlate strongly with the cut-off (lattice spacing) dependence in thermodynamic quantities. Better taste symmetry on the lattice can be achieved either by decreasing the lattice…
We present a scaling study of the QCD spectrum using a smeared P4 staggered fermion formulation, in which three, five, and seven-link staples are added to reduce the effects of flavor symmetry breaking. These studies are performed on…
LLM decoding is bottlenecked for large batches and long contexts by loading the key-value (KV) cache from high-bandwidth memory, which inflates per-token latency, while the sequential nature of decoding limits parallelism. We analyze the…
We study various improved staggered quark Dirac operators on quenched gluon backgrounds in lattice QCD generated using a Symanzik-improved gluon action. We find a clear separation of the spectrum of eigenvalues into would-be zero modes and…
Connecting multiple processors via quantum interconnect technologies could help overcome scalability issues in single-processor quantum computers. Transmission via these interconnects can be performed more efficiently using quantum…