Related papers: Development of Lattice QCD Tool Kit on Cell Broadb…
Simulations of quenched $QCD$ at relatively small but {\it nonzero} chemical potential $\mu$ on $32 \times 16^3$ lattices indicate that the nucleon screening mass decreases linearly as $\mu$ increases predicting a critical chemical…
This paper is a slightly modified and reduced version of the proposal of the {\bf apeNEXT} project, which was submitted to DESY and INFN in spring 2000. .It presents the basic motivations and ideas of a next generation lattice QCD (LQCD)…
Neural Architecture Search (NAS) has enabled automatic discovery of more efficient neural network architectures, especially for mobile and embedded vision applications. Although recent research has proposed ways of quickly estimating…
We present preliminary results for B_K, B_7^{3/2} and B_8^{3/2} from two high-statistics lattice computations. These calculations are performed at beta=6.0 and 6.2 in the quenched approximation, using mean-field-improved…
In practical satellite-based quantum key distribution (QKD) systems, the preparation and transmission of polarization-encoding photons suffer from complex environmental effects and high channel-loss. Consequently, the hinge to enhancing the…
Several possibilities exist to implement the propagation step of the lattice Boltzmann method. This paper describes common implementations which are compared according to the number of memory transfer operations they require per lattice…
Bridge++ is a general-purpose code set for a numerical simulation of lattice QCD aiming at a readable, extensible, and portable code while keeping practically high performance. The previous version of Bridge++ is implemented in double…
Transformer-based large language models (LLMs) rely heavily on intensive matrix multiplications for attention and feed-forward layers, with the Q, K, and V linear projections in the Multi-Head Self-Attention (MHA) module constituting a…
This work presents a lattice quantum chromodynamics (QCD) calculation of the nonperturbative Collins-Soper kernel, which describes the rapidity evolution of quark transverse-momentum-dependent parton distribution functions. The kernel is…
FermiQCD is a C++ library for fast development of parallel Lattice Quantum Field Theory computations. It has been developed following a top-down fully Object Oriented design approach with focus on simplicity of use. FermiQCD includes: a…
Matrix multiplication is a foundational operation in scientific computing and machine learning, yet its computational complexity makes it a significant bottleneck for large-scale applications. The shift to parallel architectures, primarily…
We investigate the computational efficiency of two stochastic based alternatives to the Sequential Propagator Method used in Lattice QCD calculations of heavy-light semileptonic form factors. In the first method, we replace the sequential…
Due to the low error tolerance of a qubit, detecting and correcting errors on it is essential for fault-tolerant quantum computing. Surface code (SC) associated with its decoding algorithm is one of the most promising quantum error…
Quantum computing has long been an experimental technology with the potential to simulate, at scale, phenomena which on classical devices would be too expensive to simulate at any but the smallest scales. Over the last several years,…
The hard mathematical problems that assure the security of our current public-key cryptography (RSA, ECC) are broken if and when a quantum computer appears rendering them ineffective for use in the quantum era. Lattice based cryptography is…
Minimising EPR consumption is the dominant objective when routing a quantum circuit on a distributed quantum computer (DQC). We present dSABRE, a SABRE-style router for multi-core processors that, on each iteration of a lookahead-driven…
Electronic devices primarily aim to offer low power consumption, high speed, and a compact area. The performance of very large-scale integration (VLSI) devices is influenced by arithmetic operations, where multiplication is a crucial…
Matrix multiplication is fundamental in the backpropagation algorithm used to train deep neural network models. Libraries like Intel's MKL or NVIDIA's cuBLAS implemented new and optimized matrix multiplication techniques that increase…
As the most central and computationally intensive component of deep neural networks, the execution efficiency of matrix multiplication directly determines the training and inference performance of models. Harnessing the parallel processing…
We calculate the leading-order matrix element for exclusive decays of $b \to s\gamma$ in the quenched approximation of lattice QCD on a $24^3\times48$ lattice at $\beta=6.2$, using an O(a)-improved fermion action. The matrix element is used…