Related papers: Two-link Staggered Quark Smearing in QUDA
We review our work done to optimize the staggered conjugate gradient (CG) algorithm in the MILC code for use with the Intel Knights Landing (KNL) architecture. KNL is the second gener- ation Intel Xeon Phi processor. It is capable of…
We discuss the CUDA approach to the simulation of pure gauge Lattice SU(2). CUDA is a hardware and software architecture developed by NVIDIA for computing on the GPU. We present an analysis and performance comparison between the GPU and CPU…
Our progress in computing the spectrum of excited baryons and mesons in lattice QCD is described. Sets of spatially-extended hadron operators with a variety of different momenta are used. A new method of stochastically estimating the…
Spin squeezing is a powerful resource for quantum metrology, and recent hardware platforms based on interacting qubits provide multiple possible architectures to generate and reverse squeezing during a sensing protocol. In this work, we…
We present and compare new types of algorithms for lattice QCD with staggered fermions in the limit of infinite gauge coupling. These algorithms are formulated on a discrete spatial lattice but with continuous Euclidean time. They make use…
We give details of our precise determination of the light quark masses m_{ud}=(m_u+m_d)/2 and m_s in 2+1 flavor QCD, with simulated pion masses down to 120 MeV, at five lattice spacings, and in large volumes. The details concern the action…
Noisy hardware forms one of the main hurdles to the realization of a near-term quantum internet. Distillation protocols allows one to overcome this noise at the cost of an increased overhead. We consider here an experimentally relevant…
Graphics Processing Units (GPUs) consisting of Streaming Multiprocessors (SMs) achieve high throughput by running a large number of threads and context switching among them to hide execution latencies. The number of thread blocks, and hence…
In this paper, we develop a new parallel auxiliary grid algebraic multigrid (AMG) method to leverage the power of graphic processing units (GPUs). In the construction of the hierarchical coarse grid, we use a simple and fixed coarsening…
Tunable couplers enable high-fidelity two-qubit gates leveraging high on/off coupling ratios and reduced crosstalk within a single design. We investigate a galvanically connected direct-current superconducting quantum interference device…
Large Language Models (LLMs) have gained popularity in recent years, driving up the demand for inference. LLM inference is composed of two phases with distinct characteristics: a compute-bound prefill phase followed by a memory-bound decode…
We present preliminary results from exploring the phase diagram of finite temperature QCD with three degenerate flavors and with two light flavors and the mass of the third held approximately at the strange quark mass. We use an order…
We simulate Quantum Chromodynamics in four Euclidean dimensions with two (degenerate mass) flavors of dynamical quarks. The Dirac operator is the so-called chirally improved operator that has been studied so far in quenched calculations. We…
Persistent homology is a crucial invariant that is used in many areas to understand data. The $O(N^4)$ run time is a hindrance to its use on most large datasets. We give a parallelization method to utilize multi-core machines and clusters.…
While lattice QCD allows for reliable results at small momentum transfers (large quark separations), perturbative QCD is restricted to large momentum transfers (small quark separations). The latter is determined up to a reference momentum…
We report on a calculation of $B_c$ ground state and radial excitation energies, obtained from heavy-charm highly improved staggered quark (HISQ) correlators computed on MILC gauge ensembles, with lattice spacings down to $a=0.044$ fm.…
We have extended our program of QCD simulations with an improved Kogut-Susskind quark action to a smaller lattice spacing, approximately 0.09 fm. Also, the simulations with a approximately 0.12 fm have been extended to smaller quark masses.…
Analysis of processing time and similarity of images generated between CPU and GPU architectures and sequential and parallel programming. For image processing a computer with AMD FX-8350 processor and an Nvidia GTX 960 Maxwell GPU was used,…
We report on progress in our study of high temperature QCD with three flavors of improved staggered quarks. Simulations are being carried out with three degenerate quarks with masses less than or equal to the strange quark mass, $m_s$, and…
We present LBcuda, a GPU accelerated version of LBsoft, our open-source MPI-based software for the simulation of multi-component colloidal flows. We describe the design principles, the optimization and the resulting performance as compared…